Practice / Incident response

Incident response and security architecture review

A real incident exposes how decisions get made under pressure. I use that same perspective when reviewing a system's design: where a failure could spread, who would notice and who could act.

Building a response function

Agree on the key responsibilities before an incident: who declares it, who directs the response, who updates the business and what evidence must be preserved. Those decisions are harder to make for the first time at 2am.

Helping create the Netflix SIRT shaped most of how I think about this. The wider security organization was 4 engineers when I joined Netflix and grew past 100, and SIRT was built inside that. What I took from it is that most of what makes response work is organizational rather than technical. When things fall apart under pressure the tooling is rarely the reason. It's usually a handoff nobody had rehearsed.

  • Incident response program design and readiness
  • Reviewing an incident after the fact, without turning it into blame
  • Security engineering and architecture review
  • Detection and escalation paths that match how the team actually works

What an architecture review should find

A good review isn't an inventory of everything imperfect, because anyone can produce that. It should name the few places where one failure crosses a boundary it never should: an identity with too much reach, a trust assumption that made sense two systems ago, a path where an incident stays invisible until a customer tells you about it.

Then it has to be short enough that a team can act on it this quarter and specific enough that nobody has to guess what changes.

Read the original writing

Ongoing long-form writing on leadership, cybersecurity and decision-making, published openly. Read First Order (opens in a new tab)

Work with me on thisBackgroundWriting & Publications