6 MIN READ
Cloud incident response runs on a different surface from the on premises version. Console, data sources, legal exposure, all of them change shape. The responder who treats a cloud incident like a 2018 data centre breach loses the first 24 hours, which is when the most important decisions get made. Here is how cloud IR actually works in 2026, and what the platform org needs in place before the page goes out.
What makes cloud IR different
Three things separate the cloud version from the on premises version. None of them are subtle.
The speed of the data. CloudTrail, Azure Activity Log, and GCP Cloud Audit Logs record every API call in near real time. A single query against them returns answers in seconds. The on premises equivalent (Windows Event Log forwarding, syslog) records events with minutes to hours of latency. The cloud responder can answer what happened in the last five minutes from a laptop. The on premises responder often cannot.
The scope of the data. A single cloud account produces millions of events per day. Most of them are noise. The job is to know which ones matter, have the queries ready, and have the dashboard open before the incident hits. Build the playbook at peace. Use it at war.
The legal exposure. The cloud provider’s terms of service impose a cooperation obligation on the customer. A law enforcement subpoena hits the provider before it hits the customer. If the logs were not preserved in the right way, the evidence is already gone by the time anyone reads the alert.
The cloud IR playbook
Five phases, roughly aligned with the NIST IR framework, adapted for cloud. Each phase assumes the platform org has done the preparation in phase one.
1. Preparation. Centralise the cloud audit logs (CloudTrail Lake, Azure Sentinel, GCP Chronicle). Document the IAM baseline. Write the runbooks for the top ten incident types. Have a working relationship with the cloud provider support team before the page goes out. None of this is exciting. All of it is what gets the responder through the first hour without fumbling.
2. Detection and analysis. The cloud native detection tools (GuardDuty, Security Command Center, Microsoft Defender for Cloud) do the heavy lifting on suspicious activity. The SIEM correlates the cloud events with the endpoint, identity, and network events. Pre built queries for the top ten incident types collapse the scope of the incident inside the first thirty minutes.
3. Containment. Isolate the affected resources: compromised IAM credentials, EC2 instances, S3 buckets. The cloud native tools (IAM policy deny, security group update, instance stop, bucket policy update) do the job from a laptop, with no physical access, no ticket, no waiting on a facilities team. Containment in the cloud is faster than it was on premises, mostly because the consoles are.
4. Eradication. Remove the attacker’s foothold: the backdoor account, the persistent IAM role, the modified Lambda function, the cron job the attacker scheduled. The cloud audit log gives a complete list of the changes the attacker made, which means the eradication work is thorough. It is also auditable later, which is the part most responders skip when they are tired.
5. Recovery and post incident. Restore the affected systems from clean backups. Rotate every credential the attacker could have touched. Write the post incident report. Update the runbook. The first three restore normal service. The last one makes the next incident shorter.
What the runbook should cover
The top ten incident types the cloud responder is going to face, in roughly that order of frequency. Each one is a runbook worth writing before the incident, not during.
1. Compromised IAM credentials. The access key, the long lived credential, the SSO credential, the federated identity. The runbook covers credential rotation, session invalidation, policy review, and a full activity review for the credential’s lifetime.
2. Compromised EC2 instance. The web server, the application server, the database server, the jump host. Isolate the instance, snapshot for forensics, rotate the credentials it touched, and rebuild from a clean image rather than trying to clean the live one.
3. Compromised S3 bucket. The data bucket, the backup bucket, the log bucket, the static asset bucket. Review the access policy and the access log, classify what was actually exposed, and restore from the immutable backup. Most of the damage in a bucket compromise is the exposure, not the loss.
4. Compromised Lambda function. The data processing function, the API trigger function, the scheduled function, the event driven function. Review the function code, review the execution log, rotate the IAM role, redeploy from a known clean version.
5. Compromised container. The Kubernetes pod, the Docker container, the serverless container, the Fargate task. Isolate the pod, review the image, review the runtime, redeploy from a known clean image.
6. Compromised database. The RDS instance, the DynamoDB table, the Aurora cluster, the document database. Review the access, review the queries, review the data, rotate the credentials.
7. Compromised network. The VPC, the subnet, the security group, the network ACL. Review the rules, review the flow logs, review the traffic, reset the rules to a known clean baseline.
8. Compromised KMS key. The encryption key, the data key, the CMK, the customer managed key. Rotate the key, review the key policy, review the usage, re encrypt the data the key protected.
9. Compromised secret. The Secrets Manager secret, the SSM parameter, the HashiCorp Vault secret. Rotate the secret, review the access, review the usage, clean up the leaked secret everywhere it landed.
10. Compromised billing. The compromised account used to mine cryptocurrency, to send spam, to run a denial of service. Page on the billing alert, review the usage, terminate the resources, lock the account down before the bill lands.
What to do this quarter
Build the centralised log pipeline. CloudTrail, Azure Activity Log, GCP Cloud Audit Logs all flowing into one place the team can query. This is work that has to be done before the incident, not during it. The platform org that built the pipeline at 2 PM on a Tuesday will find the lateral movement in the first five minutes. The one without the pipeline finds it in the post incident report, two weeks later, when the lessons are too late to be useful.
Write the runbook for the top three incident types. Compromised IAM credentials, compromised EC2 instance, compromised S3 bucket. These three cover the majority of the cloud incidents a typical platform org will see in a year. A runbook that has been written, tested, and updated is the difference between a one hour incident and a one week one.
Run a tabletop exercise against the cloud environment. The gaps in the on premises tabletop do not show up in the cloud version. The cloud specific failure modes (IAM policy mistakes, key compromise, bucket misconfigurations) need their own exercise, their own failure list, and their own follow up actions.

The bottom line
Cloud IR is faster than on premises IR, and the failure modes are different. The platform org with centralised logs, written runbooks, and a tested tabletop handles the cloud incident in hours rather than weeks. The one without those things ends up writing the runbook during the incident, in public, with legal on the call.
Sources & Further Reading
All claims in this article are sourced from primary documentation, vendor advisories, and reputable security researchers.
Spotted an error? Email the editor. Corrections are issued with a visible correction note.
Editorial standards. Every article on humanrequired.org is reviewed by a human editor before publication. AI may assist with drafting or research; final editorial control is human. Read the full standards.



