Topic 7.5
Incident Response on AWS: The First Hour
In one line
A security incident on AWS usually starts with a GuardDuty finding, an unexpected bill, or a leaked key report. The response follows a plan prepared in advance: confirm and scope with CloudTrail, contain (deactivate keys, revoke role sessions, isolate instances with a quarantine security group), preserve evidence (snapshots, logs) before changing things, eradicate the cause, recover from known-good sources, and write a blameless review. Preparation, from playbooks to break-glass access and practice, decides how well the first hour goes.
Think of it like this
A fire drill. When the alarm goes, nobody invents a plan: you know the exits, who calls the fire brigade, and where to meet. You close doors to stop the spread (contain), don't throw away the burnt toaster the inspector will want (evidence), and only then clean up and reopen.
Words you'll meet
New words in this topic, in plain English. Come back here whenever one feels fuzzy.
- Incident response
- The planned process for detecting, containing, and recovering from security incidents.
- Containment
- Steps that stop an incident from spreading or causing more damage.
- Quarantine security group
- A security group with no rules, used to cut an instance off from the network.
- Forensic snapshot
- A copy of a disk taken to preserve evidence before changes are made.
- Session revocation
- Invalidating temporary credentials already issued for a role.
- Playbook
- A written, step-by-step guide for responding to one type of incident.
Step by step
01Detect: an alert at 02:14
GuardDuty raises a high-severity finding: an access key belonging to the old CI user is being used from an unfamiliar IP address to launch instances. The on-call security engineer (Priya this week) opens the playbook for leaked credentials.
02Scope: what did the key do?
Before changing anything big, Priya checks what the key did across all regions. CloudTrail shows instances launched in an unused region and an attempt to create a new user (a common way attackers keep access).
03Contain and preserve
Priya deactivates the key (deactivate, not delete, so it stays visible in the investigation), snapshots the unknown instances for evidence, isolates them, then stops them. A second engineer joins as scribe, recording every action with timestamps.
Break it on purpose
Errors are the best teachers. Make each change, read the error, guess what went wrong, then reveal the answer.
Break #1
Terminating first, asking later
In an earlier incident, a panicked engineer terminated a suspicious instance right away and deleted the key.
Myth vs fact
Myth
Deleting a leaked access key ends the incident.
Fact
The attacker may already have created other users, keys, roles, or backdoors, or obtained temporary credentials still valid for hours. Scope everything the key did, revoke sessions, and remove anything it created.
Pro corner
Extra depth for experienced readers. New to this? Skip it for now and come back later.
- ▸
Automate the first containment steps for clear-cut findings (for example, a Lambda triggered by EventBridge on specific GuardDuty finding types that deactivates the key and notifies the channel), but keep a human in the loop for anything that affects production availability.
Remember this
- 1
Phases (NIST style): prepare → detect and analyse → contain → eradicate → recover → learn. Most damage is limited or made worse in the containment step.
- 2
Common AWS incidents: leaked access keys (pushed to a public repo), crypto-mining on hijacked compute, public S3 buckets, compromised instance credentials via SSRF to the metadata service, an over-privileged CI role misused.
- 3
Scope with CloudTrail: what did this key or role do, from which IPs, in which regions, since when? Look for
CreateUser,CreateAccessKey,AttachUserPolicy,RunInstancesin unusual regions,PutBucketPolicy,GetSecretValue. - 4
Contain credentials: deactivate leaked access keys, and for roles, revoke active sessions (an IAM policy denying tokens issued before now, which the console's 'Revoke active sessions' adds). Temporary credentials stay valid until expiry otherwise.
- 5
Contain instances, preserve evidence: snapshot EBS volumes and capture metadata first, then move the instance to a quarantine security group with no rules, detach it from the ASG and load balancer, and tag it. Don't just terminate it: you'd lose evidence.
- 6
Prepare: playbooks per incident type, a break-glass role, the security account with read access everywhere, GuardDuty and Security Hub on everywhere, contact list including AWS Support, and practice with game days.
Explain it without notes
Why preserve evidence before cleaning up?
Why isn't deactivating a role's policies enough to stop an attacker using stolen temporary credentials?
Practice
Write a one-page playbook for 'access key leaked to a public GitHub repo'.
Practise isolating an instance in a sandbox without losing evidence.
Trade-offs
- ↔
Automatic containment is fast but can disrupt production on false positives. Manual steps are careful but slow at 2 a.m. Keeping instances running preserves memory evidence but continues cost and risk.
Done when you can
Playbooks exist for leaked keys, compromised instances, and public data.
I scope with CloudTrail across all regions before cleaning up.
I contain credentials and instances while preserving evidence.
We practise and hold blameless reviews.