Forge University

Practiced, Not Improvised: Designing Response Plans and Pre-Incident Readiness

Task 2.1 tests preparation, not reaction. Before any alarm fires, a candidate needs documented, repeatable response procedures and services already configured to support them -- because the moment an incident actually starts is the worst possible time to be improvising access, tooling, or protections for the first time.

Runbooks: Turning Response Into a Repeatable Procedure

A runbook is a documented, step-by-step response procedure for a specific incident type (compromised credentials, ransomware on an EC2 fleet, an exposed S3 bucket) so that response doesn't depend on someone improvising under pressure. Systems Manager OpsCenter centralizes operational issues -- including security findings routed to it from GuardDuty or Security Hub -- into a single place where a responder can see context, run pre-built Automation runbooks, and track resolution to closure, turning a raw finding into a tracked, actionable task rather than an alert that gets lost in a channel. Amazon SageMaker AI notebooks appear on this task as a documented environment for building and running custom investigative or automated-response logic when off-the-shelf tooling isn't sufficient -- a less common exam answer than OpsCenter, but tested as an option for bespoke response tooling that a team maintains itself.

Configuring Services to Be Ready Before an Incident

Runbooks alone are not readiness -- the services and access a runbook calls on have to already be configured, tested, and waiting. That includes provisioning access ahead of time: a break-glass IAM role scoped to exactly what a responder will need, created and validated before it's needed under pressure, not improvised in the middle of an incident. It includes deploying security tooling proactively -- GuardDuty, Security Hub, and log aggregation already running continuously, not switched on in reaction to a suspected compromise, since a detection service can only find what it was watching for before the event. It includes minimizing blast radius through account and network segmentation, so a compromise in one workload account can't trivially reach every other account in the organization. And it includes configuring AWS Shield Advanced protections in advance on internet-facing resources that would be high-value DDoS targets, since Shield Advanced's stronger protections and AWS Shield Response Team engagement require pre-configuration -- there is no meaningful way to "turn on" that level of protection mid-attack and get the same benefit.

Key Mechanics

  • A runbook is a documented, repeatable procedure for a specific incident type, removing improvisation from the moment response actually matters.
  • Systems Manager OpsCenter centralizes findings (including from GuardDuty/Security Hub) with runnable Automation runbooks and resolution tracking.
  • SageMaker AI notebooks serve as an environment for custom investigative or response tooling when pre-built options aren't enough.
  • Break-glass access must be provisioned and tested in advance, not improvised during a live incident.
  • Detection tooling (GuardDuty, Security Hub, log aggregation) must already be running before an incident -- it can't retroactively see what it wasn't watching.
  • Shield Advanced protections require pre-configuration on the protected resource; there's no equivalent benefit to activating it mid-attack.

Exam Tip: Any answer describing provisioning access, deploying tooling, or configuring Shield Advanced "during" or "in response to" an incident is almost always the wrong choice on this task -- the exam consistently rewards pre-incident configuration.

Exam Tip: "Centralizes findings from GuardDuty and Security Hub with runnable automation and resolution tracking" is Systems Manager OpsCenter's signature description -- distinguish it from Security Hub itself, which generates and aggregates the findings OpsCenter tracks.

Exam Tip: Blast-radius minimization through account/network segmentation is a design-time decision tested here as incident-response preparation, not just a general security-architecture best practice -- connect it back to limiting how far a compromise can spread before containment even begins.

Worked example: A security team wants to be ready to respond to a credential-compromise incident without scrambling to figure out access or tooling mid-crisis. In advance, they create and test a scoped break-glass IAM role with exactly the permissions an incident responder would need, document the response steps as a Systems Manager Automation runbook tracked through OpsCenter, and confirm GuardDuty and Security Hub are already running continuously across the organization -- so that when a real credential-compromise finding does appear, the team executes a rehearsed procedure with pre-provisioned access rather than inventing one under pressure.

Knowledge check

Click an option to check yourself — this is a self-check, not graded or saved. The graded version pooling this module's questions is on the syllabus page.

1. A security team is preparing for potential incidents on internet-facing, high-value production resources. Which of the following must be configured in ADVANCE of an incident to provide its intended benefit, rather than activated during one?

2. A team wants security findings from GuardDuty and Security Hub to land in a single place where a responder can see full context, run pre-built automation, and track the issue through to resolution, rather than as an alert that disappears into a chat channel. Which AWS capability is designed for this?

3. An incident response plan calls for a scoped IAM role with elevated permissions that responders would use only during a confirmed security incident. When should this role be created and tested?

Log in to chat with your AI Mentor about this lesson.