Cyber Resiliency Board Briefing 12: The Recovery Decision Problem: Why Crisis Governance Must Operate at Machine Speed Without Losing Human Accountability

Executive Summary

The decisions that determine how badly a cyber incident ends are now taken faster than any governance committee can convene. Attack tempo has compressed into minutes or hours, and response and recovery tooling has been automated to keep pace. Governance has not moved at all. The practical consequence is that the most consequential decisions in a crisis get taken by whoever is awake, or by a system acting on authority nobody remembers granting. The Board's lever is to settle in advance, and in daylight, which decisions may execute automatically, which require a named individual with standing authority, and which the Board keeps for itself.

The Tempo Mismatch

CrowdStrike's 2026 threat report puts the average breakout time, the gap between an intruder landing and moving laterally, at 29 minutes. The fastest observed was 27 seconds. In one intrusion, exfiltration began within four minutes of initial access.

Set that against how a crisis decision actually gets made. Somebody notices. Somebody escalates. A bridge call is convened, participants are found and briefed, counsel is consulted, and a decision emerges. On a good day, that is two hours. On a Sunday in August, it is considerably longer. The attacker has finished and gone home before the second agenda item.

Asking people to move faster will not close this gap, but moving the decision earlier will.

Borrowed From the Exchanges

Stock exchanges solved a version of this in the 1980s. When an index falls through a defined threshold, trading halts automatically. No committee sits. Nobody is telephoned. The judgement was made years earlier by people who thought carefully about it while nothing was on fire, then encoded so that it executes without them.

The accountability did not disappear into the machinery. It sits with the people who set the threshold, documented, reviewed, and revisited. That is the model worth importing.

Three Classes of Decision

Most recovery decisions fall into one of three categories, and the useful exercise is sorting them.

Some must execute automatically, because no human is fast enough to help. Suspending a credential being used to move laterally. Halting replication so that corruption does not propagate into the copies you are retaining to recover from. These need defined triggers, defined limits, and a named owner who reviews them quarterly.

Some need one person with standing authority and no obligation to consult. Declaring a major incident. Invoking recovery. Severing a connection to a third party. What goes wrong here is usually that nobody at three in the morning believes the call is theirs to make.

Some belong to the Board and cannot be delegated anywhere. Paying an extortion demand. Suspending a regulated service. Accepting degraded operation for a sustained period. Speaking publicly. For each of these, the Board needs a rehearsed route to a quorum inside an hour, weekends included.

Automation Is Not a Blank Cheque

Automated containment destroys availability as efficiently as it preserves integrity. An isolation rule that fires during a payroll run, or a replication halt that stops the business from taking orders, can cost more than the intrusion it interrupted.

Pre-authorisation without boundaries is a delegation that the Board will come to regret. Every automated action needs a stated blast radius, a maximum duration before a human must confirm it, and a documented answer to what happens when it fires on a false positive at the worst possible moment.

Reconstructing Accountability Afterwards

Regulators, insurers, and litigants eventually ask the same four questions about every significant action taken during an incident. Who decided. On what information? Under what authority? At what time?

Where the action was automatic, the accountable human is whoever approved the policy, which is why that approval needs a name and a date on it. Organisations that assemble this record during the crisis produce a narrative. Organisations that assemble it beforehand produce evidence.

Obligations here vary by jurisdiction and sector, and it is ground to walk with counsel rather than alone.

What Good Looks Like

Cyber resilient organisations treat decision rights as part of the recovery architecture rather than as an annex to the incident response plan. They classify recovery actions by the speed at which they must be taken, and they assign each class a different authority model. Actions that must be executed without a human have documented triggers, stated limits on scope and duration, a named approver with a date against the approval, and a scheduled review. Actions requiring a single decision-maker have a named role, a deputy, and a delegation that survives holidays and resignations.

They define the small set of decisions the Board retains, and they rehearse reaching a quorum for them under realistic conditions rather than assuming it can be done. They test the delegation itself during exercises, placing participants under time pressure with incomplete information, and they treat hesitation over who owns a call as a finding rather than an inconvenience.

They record decisions as they are taken, capturing authority, trigger, information available, and timestamp, using a method that works when the corporate collaboration tooling is unavailable. They review automated actions after every exercise and every incident, retiring triggers that no longer reflect the estate and tightening those that fired too broadly.

Most importantly, they treat the granting of authority as a deliberate governance act performed in advance, and the crisis itself as the point at which that authority is exercised rather than invented.

Questions to Ask Your Executive Team

  • Which recovery actions execute today with no human in the loop, and who approved that, by name and date?

  • What is each automated action forbidden to do, and who can countermand it while the incident is running?

  • Who can declare a major incident and invoke recovery at three in the morning without consulting anyone, and have they ever done so in an exercise?

  • How long does it take us to assemble a quorum of directors at a weekend, measured rather than estimated?

  • Where is the decision record kept, and does it survive the loss of our normal collaboration tooling?

  • If asked in two years who made a particular call during an incident, could we answer with a record rather than a recollection?

Key Takeaways for the Board

  • Attack tempo is measured in minutes; governance tempo is measured in hours. The gap closes by moving decisions earlier, not by moving people faster.

  • Every automated response is a standing delegation of Board authority, whether or not it was ever described that way.

  • Pre-authorisation without stated limits transfers risk to the Board rather than managing it.

  • Accountability cannot be delegated to a script. It attaches to whoever approved the policy the script executes.

The Board's Role

The Board does not need to decide which playbook step fires automatically or what threshold triggers an isolation rule. It should, however, expect management to have a coherent model for how authority is granted, bounded, and recorded before a crisis begins. Management should be able to describe which recovery actions run without human involvement and on what triggers, what each of those actions is forbidden to do, who holds standing authority to declare an incident and invoke recovery, and which decisions have been reserved to the Board. They should be able to explain how delegations survive absence and turnover, how decisions are recorded while systems are down, and how automated triggers are reviewed after exercises and incidents.

The Board should also challenge response metrics that stop at speed. A dashboard reporting Mean time to Detect and Mean time to Respond says very little about whether the actions taken were ones anyone had the authority to take, whether the organisation could demonstrate that authority afterwards, or whether recovery was to a trusted state. A more useful view comes from management showing which decision classes have been mapped, which delegations have been exercised under time pressure, and where the gaps remain.

The fundamental Board-level question is this: "Can management demonstrate that every action taken in the first hour of a crisis would be one somebody was authorised to take, and that we could prove it?"

Closing Thought

Machine speed and human accountability sit together comfortably, provided the human judgement happens first. The Board will not be in the room for the decisions that matter most, because those decisions get taken in minutes by people and systems acting on whatever authority was granted beforehand. Granting that authority deliberately, with limits and a record, is the governance work, and the only time it can be done well is while the building is quiet.

Next Briefing

Cyber Resiliency Board Briefing 13: The Rehearsal Problem - why most cyber exercises confirm what the organisation already believes, and what it takes to design one that tests delegated authority rather than technical recovery.

Previous
Previous

17,600 Actions and Nothing to Go On: What Agentic Intrusions Do to Your Investigation

Next
Next

Cyber Resiliency Board Briefing 11: The Reconnection Problem: Why Bringing Systems Back Can Be More Dangerous Than Restoring Them