Cyber Resiliency Board Briefing 11: The Reconnection Problem: Why Bringing Systems Back Can Be More Dangerous Than Restoring Them
Executive Summary
After a destructive cyberattack enormous attention is placed on restoration.
Can the backup be recovered? Has the server been rebuilt? Has the application been tested? Have persistence mechanisms been removed? Have credentials been rotated? Has the vulnerability used by the attacker been remediated?
Eventually, however, every recovered workload reaches another critical decision gate: can this service be returned to production?
That does not necessarily mean reconnecting it to the environment from which it came.
In a destructive cyberattack, I generally recommend thinking in terms of three distinct environments:
The dirty or unknown environment, containing systems that are known to be compromised or whose trust state has not yet been established.
The clean room, where systems and data are investigated, root cause and persistence are understood, attack surface is assessed, and remediation takes place through cleaning, patching, rebuilding or recovery from an appropriate trusted point.
The clean production environment, into which systems that have passed the required assurance gates are repatriated and from which recovered business services can progressively resume.
Systems that are subsequently proven not to have been impacted can also be moved from the dirty environment into clean production, subject to appropriate validation.
This architecture changes the recovery objective as rather than trying to make the compromised environment trustworthy again while simultaneously restoring services into it, the organisation creates a progressively expanding area of known-good production.
Where business dependencies still exist in the dirty or unknown environment, controlled connections can be established from clean production back into it. Those connections should be treated as exceptions, constrained to what is required, monitored closely and removed as remaining dependencies are recovered.
This significantly reduces the risk of taking a thoroughly investigated and remediated workload and placing it straight back into an environment whose overall trust state remains uncertain.
The Board should therefore ask their operational teams: “Do we have a recovery architecture that allows us to progressively rebuild trusted production without making recovered systems dependent upon infrastructure we still do not trust?”
Recovery Does Not End When the Workload Starts
Traditional disaster recovery tends to have a relatively simple end state: The server is restored. The application starts. Users can access it. The service returns to production.
A destructive cyberattack creates a much more complicated problem because the recovery team is operating in an environment where trust has been deliberately undermined.
Before a recovered workload can safely return to production, the organisation needs confidence in:
The recovery point selected.
The integrity of the operating system, application binaries and libraries.
Configurations.
Local and service accounts.
Credentials.
Certificates, API tokens and secrets.
Dependencies.
Network paths.
Identity services.
Security controls.
Management interfaces.
Connected third parties.
Data integrity.
The systems with which the workload will communicate.
This is why I recommend separating recovery into distinct trust zones:
The clean room exists to establish sufficient confidence in the workload itself.
The clean production environment provides somewhere that has met the organisation’s defined assurance threshold for that workload to go once sufficient confidence has been established.
The dirty environment remains contained until its individual components can either be validated and migrated, remediated through the clean room, or retired.
That separation prevents the organisation from turning every recovery decision into a bet on the trustworthiness of the entire existing estate.
Three Environments, Three Different Purposes
It is useful to think of the recovery architecture as three connected but deliberately separated environments.
The Dirty or Unknown Environment
This is the environment inherited from the incident. Some systems may be demonstrably compromised and others may simply have an unknown state.
It is important to remember absence of evidence of compromise should not automatically be interpreted as evidence that a system is safe.
The dirty environment may therefore contain a mixture of:
Confirmed compromised systems.
Systems still awaiting investigation.
Systems believed to be unaffected but not yet sufficiently validated.
Legacy infrastructure that cannot immediately be replaced.
Dependencies required by business services during recovery.
Third-party connections whose trust state is still being established.
The key point is that this environment does not automatically regain trust because the attacker appears to have been contained.
The Clean Room
The clean room is where evidence is converted into confidence.
Systems entering the clean room are investigated to determine:
Root cause.
Initial access.
Attacker persistence.
Lateral movement.
Compromised credentials.
Malicious tooling.
Exploited vulnerabilities.
Configuration weaknesses.
Excessive privilege.
Unnecessary services or exposed attack surface.
The outcome of the investigation may require volume-based recovery and remediation, but often the faster or more efficient answer is to rebuild from trusted vaulted install images and configurations, then bring production data back from backups.
The clean room therefore performs two related functions:
It helps the organisation understand why the system became compromised.
It provides an environment in which the conditions that enabled, or are the result of, that compromise to be addressed before the workload is returned to service.
The Clean Production Environment
The clean production environment is the destination. It is here that the systems that have passed the required investigation, remediation, validation and business testing gates are repatriated into it.
It should contain trusted foundational services such as:
Identity.
DNS.
Core networking.
Privileged access.
Security monitoring.
Endpoint protection.
Management infrastructure.
Required communications services.
Additional applications and business services are added in controlled recovery waves, creating a progressively expanding trusted estate.
Rather than attempting to clean an entire enterprise in place, the organisation is effectively building a new area of trusted production and migrating towards it.
Build the Clean Island Before Reconnecting the Continent
A useful way of thinking about this is as a clean island.
Immediately after the incident, much of the existing enterprise may represent uncertain territory.
The objective is to establish a small, defensible area of known-good infrastructure first. That clean island might initially contain:
Recovered identity.
DNS.
Core network services.
Security tooling.
Administrative services.
Communications.
A handful of Minimum Viable Company services.
As assurance increases, more systems can move onto the island.
Eventually the clean environment becomes the normal production environment, while the dirty estate contracts until it can be decommissioned.
This is safer than attempting to prove that every system, identity, network route and management interface across the existing estate is clean before recovery can progress. It also allows the organisations to bring services back in a tiered manner, starting with the Minimum Viable Company services first.
Most importantly, it provides an architectural mechanism for managing uncertainty.
You do not need to prove the entire enterprise is trustworthy before restarting business operations, you just need enough trusted infrastructure to support the next recovery wave.
Controlled Connections Back Into Dirty
Operational reality means the clean production environment cannot always be entirely self-contained from day one, and this should be planned for in advance.
Complete separation will not always be immediately achievable. Critical business services may still depend on legacy databases, industrial controllers, third-party connections or application dependencies that cannot yet be rebuilt, migrated or redesigned.
In those cases, the objective should be to reduce and govern the residual risk rather than allow it to dictate the recovery model. Any connection from clean production back into the dirty or unknown environment should therefore be treated as an explicit exception: narrowly scoped, technically constrained, closely monitored, time-limited and subject to senior risk acceptance.
The objective is to prevent temporary recovery connectivity from quietly rebuilding the same broad trust relationships that existed before the attack. Every bridge from clean into dirty should therefore be treated as a potential route back across the trust boundary.
Reconnection Rebuilds the Attack Surface
During the containment phase of a major incident, organisations will deliberately destroy connectivity to restrict command annd control, stop further data exfiltration, constrain attacker movement and reduce the available attack surface. Networks may be isolated, VPNs disabled, federation relationships suspended, privileged accounts locked, APIs disconnected and third-party access blocked. Administrative interfaces can be heavily restricted, while entire data centres or cloud environments may effectively be quarantined.
Recovery gradually reverses this process. As connectivity is restored, the operating environment expands again and with it the number of trust relationships, authentication paths and reachable systems. Re-enabling APIs allows systems to exchange information, restoring firewall rules increases reachability, and returning administrative platforms to service rebuilds management capability while potentially recreating opportunities for an attacker.
The objective should therefore be to restore the connectivity required to operate the business without automatically recreating the previous network model. Recovery provides an opportunity to reconsider which connections genuinely need to exist and which were simply inherited through years of architectural change.
Where possible, the clean production environment should be designed around minimum necessary connectivity. Firewall rules, legacy trust relationships and dormant administrative paths should be validated against current business requirements before they are reinstated, rather than copied wholesale from the pre-incident environment.
A major cyberattack is one of the few occasions when organisations have both the mandate and the opportunity to challenge accumulated connectivity. Recovery can therefore do more than restore operations: it can leave the organisation with a materially smaller attack surface at a point when executive attention, funding and organisational appetite for change are often at their highest.
Repatriation Is the Critical Transition
The most important transition in this model is from the clean room into clean production. Crossing that boundary requires a deliberate judgement that sufficient evidence has been gathered to allow the workload to participate in a trusted production environment.
The evidence required will vary according to the workload, its criticality and the nature of the incident. It may include selection of an appropriate recovery point, forensic investigation, understanding of root cause, removal of persistence mechanisms, remediation of the attack vector, vulnerability patching, credential rotation, replacement of certificates or secrets, reduction of unnecessary attack surface, validation of application configuration, malware scanning, technical testing, business-service validation and deployment of the required security tooling. Dependencies must also have been assessed to determine whether they are trusted, isolated appropriately or subject to an explicitly accepted residual risk.
Repatriation should therefore operate as a formal recovery gate, with evidence, ownership and risk acceptance proportionate to the importance of the service being returned to production.
Systems From Dirty May Also Move Directly Into Clean Production
Not every system in the original environment needs to pass through an identical recovery process. Some systems, when investigated, may show that they have not been affected. Those systems can potentially be migrated from the dirty or unknown environment directly into clean production.
This decision should be evidence-led and based on pre-established criteria for determining when a system has enough assurance to make that move. Depending upon the incident, those criteria might include:
No evidence of attacker interaction.
Appropriate endpoint and forensic telemetry.
Validation against known attacker TTPs.
Credential rotation.
Patch and vulnerability status.
Security-control validation.
Review of relevant logs.
Configuration assurance.
Business testing.
The important point is that location does not determine trust.
A system is not clean because it escaped obvious encryption: persistence beachheads are often left unencrypted. Likewise, a system does not necessarily require rebuilding simply because it existed inside the affected estate. Much like in the medical profession, evidence should always drive treatment.
Recovery Waves Need Trust Boundaries
This becomes particularly important as recovery moves beyond the Minimum Viable Company.
The earliest recovery wave usually contains a relatively constrained collection of essential capabilities:
Core identity.
DNS and network services.
Critical security tooling.
Essential management infrastructure.
Priority business services.
Critical communications.
Required regulatory, safety or operational capabilities.
These systems establish the initial clean production environment, with additional recovery waves expanding from it.
Each wave should have:
Defined entry criteria.
Required security controls.
Known dependencies.
Evidence requirements.
Validation procedures.
Repatriation criteria.
Defined decision authority.
Monitoring requirements.
Rollback procedures.
The objective is to prevent recovery becoming a race in which systems are moved simply because they are available: availability for repatriation should not determine recovery order.
The Minimum Viable Company and agreed business priorities should determine recovery order, while evidence and risk appetite determine whether a workload is ready to cross into clean production.
Repatriation Is a Risk Decision, Not a Technical One
One of the recurring themes across this Board Briefing series, and my advisory services to customers is that cyber recovery is full of decisions that look technical but are, in substance, business risk decisions. Repatriation is a good example.
By the time a system is ready to return, a great many people will have done their part. Security will have gathered evidence. Infrastructure will have confirmed the rebuild, application teams will have finished functional testing, and incident responders will have given their view on the likelihood of persistence. Threat intelligence may have added what is known about the attacker’s behaviour, identity teams will have rotated the relevant credentials, and the business will have made clear, probably more than once, that the service is needed urgently.
None of that answers the question that actually matters: “is the evidence in hand sufficient to allow this system into clean production?”
That question rarely comes with a tidy answer. Complete forensic evidence may never have been available. A legacy application might not support one of the new security controls. Some dependencies will have been validated more thoroughly than others, and a critical workload may temporarily need connectivity back into a part of the estate that has not yet been assessed. The decision may still, quite reasonably, be to proceed..
The Recovery Factory Needs Quality Gates
At enterprise scale, systems need to move through a repeatable pipeline of investigation, remediation, restoration, validation, testing and approval. Throughput matters because hundreds or thousands of workloads may need to progress through that pipeline simultaneously.
The workflows and supporting technology I recommend in my consulting engagements are designed to build recovery as a factory rather than as a heroic effort. The process is defined in advance, automated wherever a step is well understood, and produces the same result whether it is the first system through the line or the five hundredth. That matters because the constraint in a real recovery is never the technology; it is people, time and attention. An organisation with a handful of engineers who know the estate cannot afford to have them improvising under pressure across hundreds of systems. The factory model turns that scarce expertise into design effort spent once, up front, and lets the production line carry the load when it counts.
Factories do not optimise purely for throughput. They also have quality control, because increasing production while removing inspection simply lets defective products leave the building faster.
Cyber recovery carries the same risk. As the pressure to increase throughput builds, assurance tends to weaken in ways nobody actually decided on. Investigations get shorter and validation gets compressed. Exceptions that were once remarkable become routine, repatriation approvals slide from formal sign-off into a nod on a call, and an overwhelmed security team starts producing evidence that varies from one system to the next. Before long, teams are assuming somebody else performed the check, and business pressure is quietly doing the work that technical confidence should be doing.
The recovery factory therefore needs explicit quality gates. Depending on the organisation and the workload, those gates will typically require evidence that an appropriate recovery point has been selected, that the relevant forensic investigation is complete and the root cause and persistence mechanisms are understood, and that the attack surface which enabled the compromise has been addressed along with the relevant vulnerabilities. They will want confirmation that credentials, certificates and secrets have been rotated, that security controls are operating as intended, and that both technical and business testing have finished. Dependencies should be either trusted or reached through appropriately controlled connections, enhanced monitoring should be active, residual risk should be documented, and an authorised person should have approved the move into clean production.
The exact controls will vary considerably between organisations, but the principle should not: nothing enters clean production simply because somebody says it is ready. There should be evidence.
Avoid the Binary Model of Trust
During recovery many organisations tend to talk about systems as either “trusted” or “untrusted”. Reality is rarely that convenient, because trust develops progressively. An application may have strong assurance around its operating system and binaries while doubt remains about its configuration. An identity environment may have been substantially rebuilt while a handful of legacy service accounts still await rotation. A critical business service may be running entirely inside clean production apart from one controlled dependency back into the dirty estate, and a third party may be admitted through a tightly constrained temporary connection long before normal integration is restored.
Mature recovery models therefore benefit from defining distinct trust states. A workable set is:
Dirty / Unknown: the system remains inside the affected estate and its trustworthiness has not yet been established.
Clean Room: the system is being investigated, remediated, rebuilt or validated.
Clean Production - Restricted: the workload has been admitted to clean production but runs with constrained connectivity, reduced functionality or additional monitoring because some dependencies remain unresolved.
Clean Production - Normal: the required assurance is in place and normal operating relationships can progressively resume.
What these states provide during a crisis is an alternative to the false choice between completely disconnected and completely restored. A business service can often return safely with restricted connectivity and closer oversight while recovery work continues elsewhere.
Temporary Connections Have a Habit of Becoming Permanent
Controlled connections between clean and dirty environments create a governance problem of their own. Incidents generate workarounds: temporary firewall rules, new administrator accounts, emergency VPNs, extra proxies and gateways, altered authentication requirements, and legacy integrations that linger longer than anyone intended. During recovery these measures may be entirely appropriate.
The danger comes later, once business pressure subsides and temporary arrangements quietly harden into permanent architecture. Six months on, nobody remembers why a rule allows traffic from clean production into an old network segment, or why an emergency privileged account still exists. Recovery debt has begun to accumulate.
Organisations should therefore track every temporary recovery control and cross-boundary connection with an owner, a rationale, the business dependency it supports and the risk it addresses, the date it was implemented, the conditions under which it can be removed, a review date and a formal closure. Without that discipline, the architecture designed to contain uncertainty gradually dissolves back into the architecture that existed before the incident.
Monitor Clean Production Differently
Clean production should have stronger visibility than the environment it replaces, and immediately after a destructive attack the organisation is unusually well placed to provide it. It now knows things it did not know before: how initial access was gained, what tooling and persistence mechanisms the attacker used, which identities were targeted, what command-and-control infrastructure was involved, and how privilege escalation, lateral movement, data staging and exfiltration were carried out. It knows which vulnerabilities were exploited and, usually, what the attacker was trying to achieve.
That intelligence should shape the monitoring of clean production directly. For a period, organisations should consider heightened attention on privileged authentication, newly created and service accounts, administrative tooling and remote access, scripting activity, unexpected network connections, and any reappearance of the indicators and behaviours observed during the incident. Changes to security controls, signs of persistence, data movement, traffic crossing between clean and dirty zones, and the behaviour of newly admitted workloads all deserve closer scrutiny than they received before.
If the attacker was able to operate for weeks before being detected, rebuilding exactly the same telemetry and detection capability gives little reason to expect a different outcome should they return.
Plan for Rollback
Repatriation decisions will sometimes be wrong, and that needs to be designed into the recovery model rather than discovered on the day. A workload may behave unexpectedly. Persistence that nobody had seen may surface, a credential believed to be secure may be used suspiciously, or a controlled connection into the dirty estate may reveal behaviour that changes the risk assessment. A dependency may prove less trustworthy than expected, or the organisation may learn that the attacker retained access through a mechanism it had not understood.
The response should not require improvisation. A workload admitted to clean production needs to be capable of being isolated again, cross-boundary connections need to be removable quickly, credentials may need to be disabled, and a system may need to return to the clean room for further investigation. The organisation should know who can make those calls and how fast they can be executed.
I know first-hand that the psychology of recovery pushes hard in the opposite direction. After days of disruption, nobody wants to take a recovered service away again. Resilience depends on being prepared to do exactly that when the evidence changes.
What Good Looks Like
Cyber resilient organisations establish a recovery architecture with explicit trust boundaries. They separate the dirty or unknown environment from recovery and clean production, establish a clean room for investigation, root-cause analysis, persistence identification, attack-surface assessment and remediation, and build a clean production environment rather than restoring systems back into the affected estate. Trusted foundational services are recovered into clean production first. Everything else is repatriated from the clean room only once defined assurance gates have been met, and systems proven unaffected move from the dirty estate through the same evidence-led process.
Where business dependencies require it, they permit controlled connections back into dirty or unknown environments, restricted to the minimum access necessary, closely monitored, and assigned owners and removal criteria. They recover business services in predefined waves aligned to Minimum Viable Company priorities, validating dependencies as well as individual workloads.
They require evidence before anything enters clean production, define who has authority to approve repatriation, and document residual uncertainty and risk acceptance. Where full production trust cannot yet be justified, they use restricted operating states, and they reassess legacy trust relationships rather than automatically recreating them.
They apply enhanced monitoring to clean production and use incident intelligence to improve detection within the recovered environment. They maintain the ability to rapidly re-isolate recovered services, record emergency configuration changes and compensating controls, and exercise movement between dirty, clean-room and clean-production environments during realistic recovery scenarios.
Most importantly, they treat trusted production as something that is deliberately rebuilt and progressively expanded.
Questions to Ask Your Executive Team
Do we have a defined dirty, clean-room and clean-production architecture for cyber recovery, and can we establish clean production independently of the affected production network?
Which foundational services must exist before the first business workloads can enter clean production?
What evidence must exist before a workload moves from the clean room into clean production, and what criteria allow an apparently unaffected system to move directly from the dirty estate? Who has the authority to approve those decisions?
Do we understand which workloads in clean production will temporarily need access back into dirty or unknown infrastructure, and how those connections will be constrained, authenticated and monitored?
Who owns each temporary cross-boundary connection, and what conditions cause it to be removed?
Are our recovery waves aligned to our Minimum Viable Company and agreed business priorities, and can critical services operate with restricted functionality while some dependencies remain in the dirty environment?
Would we automatically recreate the connectivity that existed before the attack, or use recovery as an opportunity to remove unnecessary attack surface?
What enhanced monitoring is applied to systems entering clean production, and how is intelligence from the incident translated into detection within that environment?
Can we immediately remove a workload from clean production if new evidence changes its trust assessment?
How do we track the emergency accounts, firewall rules, workarounds and compensating controls introduced during recovery?
During exercises, do we test migration through dirty, clean-room and clean-production environments, or does the exercise end when a backup restores successfully?
Key Takeaways for the Board
The existing production estate should be treated as dirty or unknown until sufficient evidence establishes otherwise. The clean room exists to investigate root cause, identify persistence, understand attack surface and remediate systems before they return to service, and a separate clean production environment gives recovered and validated workloads a trusted destination. Systems shown to be unaffected can also migrate from the dirty estate into clean production, provided that decision is evidence-led. As recovery continues, clean production should progressively expand while the dirty estate contracts.
Connections from clean production back into dirty or unknown infrastructure will sometimes be necessary, but they should be treated as temporary, constrained and closely monitored exceptions, and recovery should never automatically recreate the trust relationships and network paths that existed before the incident. Repatriation into clean production is a formal risk decision supported by evidence, which is why the recovery factory needs assurance gates as well as throughput.
The period immediately after recovery remains high risk, so enhanced monitoring should accompany the clean production environment, and the recovery architecture should preserve the ability to reverse decisions, remove cross-boundary connections and rapidly re-isolate workloads. Handled well, cyber recovery becomes an opportunity to rebuild production with less attack surface and stronger trust boundaries than existed before.
The Board’s Role
The Board does not need to decide which network segment a recovered server enters. It should, however, expect management to have a coherent architecture for how trusted production will be rebuilt. Management should be able to describe what constitutes the dirty or unknown environment, how the clean room is isolated and used, how clean production is established, and which foundational capabilities must be recovered first. They should be able to explain what evidence is required before systems enter clean production, how apparently unaffected systems are validated, how unavoidable dependencies on the dirty estate are controlled, and who owns repatriation and residual-risk decisions. They should also be clear on how enhanced monitoring is applied and how a workload can be removed from clean production if confidence changes.
The Board should also challenge recovery metrics that stop at restoration. A dashboard showing that 5,000 workloads have been restored says very little about whether those workloads are trusted or capable of safely delivering business services. A more useful view of organisational recovery comes from management demonstrating how many systems have moved through investigation, remediation, validation and repatriation into clean production, how many services are operating there, and what residual dependencies remain in the dirty estate.
The fundamental Board-level question is this:”can management demonstrate that we can progressively rebuild a trusted production environment while containing the systems and dependencies whose trust state remains unknown?”
Closing Thought
A destructive cyberattack changes the meaning of the production network. Before the incident, connectivity is an enabler; during recovery, that same connectivity becomes a source of uncertainty. Trying to restore thousands of systems straight back into an environment whose trust state is still being established risks mixing known-good recovery with unknown infrastructure.
A safer approach is to reverse the problem. Build a clean production environment, use the clean room to investigate, remediate and validate what will enter it, and move systems across the trust boundary as the evidence allows. Where clean production still needs something from the dirty estate, build the narrowest possible bridge and watch it closely until that dependency can be recovered as well. Over time, clean production grows and the dirty environment shrinks.
That gives recovery a direction. Rather than trying to prove that the old company is safe enough to restart, you are progressively rebuilding a trusted company alongside it, and creating somewhere trustworthy for production to return to.
Next Briefing
Cyber Resiliency Board Briefing 12: The Recovery Decision Problem: Why Crisis Governance Must Operate at Machine Speed Without Losing Human Accountability