Cyber Resiliency Board Briefing 7: Resilience of the Control Plane, The Systems That Recover the Systems
Executive Summary
Most cyber resilience programmes focus on protecting production systems, far fewer consider the resilience of the systems responsible for recovering them, regaining trust and coordination response.
When a destructive cyberattack occurs, organisations quickly discover that recovery depends upon a relatively small number of management platforms, administrative services and privileged systems. These systems orchestrate infrastructure, manage identities, control backups, automate deployments, provide monitoring and enable administrators to rebuild the enterprise.
Collectively, these can be thought of as the “recovery control plane”.
If the control plane is compromised, unavailable or untrusted, the recovery of production systems slows dramatically, or may stop altogether.
While boards will often ask their executives “Are our critical business systems protected?”, they should also be asking them “Are the systems that recover our business systems even more resilient?”
Recovery is the often called the “last line of defence” but recovery cannot begin without trusted control.
The Forgotten Critical Systems
Every organisation has systems that manage other systems. These are often invisible to the business because they rarely deliver customer-facing services directly. Examples of these include:
Active Directory and identity infrastructure
Privileged Access Management (PAM)
Backup and recovery platforms
Hypervisor and virtualisation management
Cloud management consoles
Infrastructure-as-Code repositories
Configuration management platforms
Endpoint management systems
Software deployment platforms
Certificate authorities
DNS and DHCP
Monitoring and orchestration platforms
Secrets and key management systems
These platforms are the operating system of enterprise IT itself. Without them, rebuilding becomes manual, inconsistent and significantly slower.
The Control Plane Is a High-Value Target
Modern attackers increasingly target the control plane rather than individual hosts, as compromising a management platform often provides far greater leverage than compromising a single workload. Once attackers obtain privileged administrative access, they may be able to:
Disable security controls
Delete or encrypt backups
Disrupt or eavesdrop on network communications
Push malicious configurations
Deploy ransomware at scale
Remove forensic evidence
Destroy automation.
Prevent recovery operations
In effect, compromising the control plane allows attackers to influence every system that depends upon it.
Recovery Requires Trusted Administration
One of the first questions that needs to be answered before any recovery starts is “Can we trust the systems and identities performing the restore?”. If your administrators cannot trust:
Administrative identities
Orchestration platforms
Automation scripts
Configuration repositories
Recovery tooling
then every recovery action risks reintroducing compromise or attack surface. Recovery post destructive cyberattack therefore becomes an exercise in rebuilding trust before rebuilding technology.
Separate the Recovery Plane from Production
Many organisations operate recovery tooling within the same administrative environment as production, which creates unnecessary risk. A mature cyber resilience strategy deliberately separates recovery capabilities from the production environment they protect.
Examples include:
Separate administrative identities
Independent authentication
Dedicated privileged access
Segregated management networks
Immutable recovery repositories
Independent logging and audit
Separate encryption key management
Isolated recovery environments
Not using Baseboard Management Controllers on backup systems that provide the capability to wipe storage
The objective is to prevent the failure of the systems responsible for recovery alongside the systems they recover.
Recovery Depends on a Healthy Control Plane
The ability to bring an entire organisation’s IT infrastructure back to a trusted state assumes recovery can occur repeatedly, consistently and at scale. That is only possible if the underlying control plane remains trusted.
This capability relies upon:
Trusted operating system images
Version-controlled configurations
Infrastructure-as-Code
Application deployment pipelines
Automation tooling
Identity services
Recovery orchestration
Without these components, recovery becomes largely manual, and manual recovery does not scale.
The Business Impact Is Greater Than It Appears
Boards often associate the control plane with Business-as-Usual IT operations. In the event of a destructive cyberattack its importance is much broader. If recovery cannot be coordinated:
Regulatory obligations may not be met.
Customer services remain unavailable.
Safety-critical operations may be delayed.
Financial losses continue to accumulate.
Contractual commitments cannot be fulfilled.
Executive decision-making becomes increasingly constrained.
Recovery may occur prematurely without remediating attack surface or adversary persistence, resulting in reattack and further impact.
The resilience of the control plane therefore directly affects organisational resilience.
What Good Looks Like
Cyber resilient organisations recognise the control plane as the most critical infrastructure within their own enterprise.
They:
Identify every system responsible for administering or recovering production environments.
Prioritise control plane resilience above many production workloads.
Separate recovery identities from production identities.
Maintain immutable, independently protected recovery artefacts.
Protect automation, scripts and Infrastructure-as-Code as critical recovery assets.
Regularly exercise recovery assuming elements of the control plane have been compromised.
Ensure recovery tooling operates from a separate trust boundary wherever practical.
A Practical Analogy
Imagine a fire rages in a part of a city.
The fire engines survive, the firefighters are ready but the emergency control room has been destroyed.
Dispatch systems no longer work. Radio communications have failed. Maps have been lost. Details of hazardous and flammable materials stored in buildings are unavailable. Locations of residents are unknown.
The firefighting capability still exists, but it cannot effectively be applied where it is needed most. The firefighters may waste time, effort and water on important areas, while valuable and populated buildings burn. They may move on too quickly, causing smouldering areas to reignite.
Without the control centre putting out the fire and restoring normality to the city becomes dramatically harder.
Questions to Ask Your Executive Team
Which systems constitute our cyber recovery control plane?
Are these systems more resilient than the production environments they administer?
Which control plane services represent single points of failure?
Could we recover if our primary identity infrastructure were unavailable?
Are recovery identities separated from production administration?
Have we tested recovery assuming elements of the control plane have been compromised?
How do we establish trust in our response and recovery tooling before rebuilding production?
Key Takeaways for the Board
The systems that recover the organisation are even more critical than the systems being recovered.
Modern attackers increasingly target management and administrative platforms.
Recovery depends upon trusted identities, automation and orchestration.
Recovery tooling should operate from separate trust boundaries wherever possible.
Control plane resilience should be treated as a strategic cyber resilience capability, not simply an operational IT concern.
The Board’s Role
The Board should ensure that cyber resilience programmes explicitly identify, protect and validate the control plane that enables trusted recovery.
This includes understanding which management platforms, identity systems, administrative services, investigatory tooling and automation capabilities underpin recovery, ensuring they are appropriately segregated from production, and requiring evidence that they can remain trusted during a destructive cyber event.
The Board should also challenge whether recovery plans assume the control plane survives unchanged, or whether they demonstrate how trust can be re-established if those critical systems themselves become compromised.
Closing Thought
Every organisation depends upon systems that manage other systems. When those systems fail, recovery slows..
When they cannot be trusted, recovery may stop altogether.
The resilience of your organisation ultimately depends on the resilience of the systems that recover the systems.
Next Briefing
The Third-Party Recovery Problem: Why Your Recovery Depends on Organisations Outside Your Control