Resilience by design means continuity requirements are considered before a project, system, facility, supplier or process change goes live. The objective is to prevent new recovery gaps from being introduced faster than the BCM program can discover them later.
Put resilience into the change lifecycle
Continuity review should occur at defined project gates: concept, solution design, sourcing, pre-production readiness and post-implementation review. The project does not need a full BIA at every stage, but it should identify whether the change creates or alters a critical service, dependency, recovery target, concentration risk, supplier dependency, data location, manual workaround or crisis-response requirement. Link this process with BCM project and change management.
Define resilience requirements before design is locked
Project teams need requirements that are testable. Examples include maximum tolerable outage, RTO/RPO, minimum viable capacity, offline operating needs, geographic separation, alternate supplier capability, restoration priority, data recovery, emergency access, communication fallback and recovery evidence. A requirement such as “high availability required” is too vague because it does not define the failure condition or acceptable outcome.
| Gate | Continuity question | Required evidence |
|---|---|---|
| Concept | Could this change affect a priority service? | Business impact and dependency screening |
| Design | How will the service operate or recover after plausible failures? | Recovery architecture and workaround design |
| Sourcing | Does the supplier create concentration or recovery dependency? | Continuity requirements and supplier evidence |
| Pre-go-live | Has recovery capability been demonstrated? | Test results, runbooks, owner acceptance |
| Post-go-live | Did the change create unplanned dependencies? | Updated BIA, plans and corrective actions |
Challenge single points of failure
Design review should look beyond obvious hardware redundancy. Single points can exist in people, credentials, DNS, certificates, network paths, cloud tenancy, data pipelines, specialist vendors, licensing, physical access, or a single operational procedure. A component may be technically redundant while still depending on one control plane or one team. Capture these dependencies explicitly and decide whether to remove, mitigate or formally accept them.
Use impact tolerance to test the design
Where the organization uses impact tolerances, compare the end-to-end service design with the maximum disruption the organization is prepared to tolerate. A resilient design must consider the whole service chain, not only the new component. See impact tolerance versus RTO when translating service-level outcomes into component recovery requirements.
Include degraded operation
Projects should define what happens before full restoration. Can users work from another location? Can a transaction be recorded manually? Can a secondary supplier provide a reduced service? Can a cloud application operate without one integration? Define the minimum viable service and the controls that still apply in degraded mode. Do not rely on a workaround that has never been timed or resourced.
Test before production dependency forms
Pre-go-live resilience testing should reflect the most important failure modes. Depending on the change, this may include loss of a zone or site, identity failure, database restore, supplier unavailability, network isolation, corrupted data, loss of a privileged administrator or recovery from clean infrastructure. Test evidence should include actual elapsed time, capacity, data state, unresolved defects and business acceptance.
- Confirm recovery runbooks are owned and accessible during an outage.
- Confirm monitoring can detect the failure and distinguish degraded service from total failure.
- Confirm support and supplier escalation paths work outside normal office hours.
- Confirm backup and restore evidence meets the stated data-loss tolerance.
- Confirm new dependencies are reflected in BIA, plans and service maps.
Use risk acceptance only for conscious exceptions
If the project cannot meet a continuity requirement before go-live, the exception should state the gap, scenario, business impact, interim control, accountable owner, approver, expiry date and remediation plan. “Accepted by project” is not sufficient if the risk affects a business service outside the project’s authority.
Example change review
A project replaces two regional suppliers with one global platform to reduce operating cost. Functional testing passes, but continuity review shows that all transactions now depend on the same provider and identity tenant. The resilience gate requires evidence of provider recovery, an alternate operating process, tested export capability and an approved concentration-risk decision before go-live. The purpose is not to block change; it is to make the new dependency visible and governed.
Completion criteria
A resilience-by-design review is complete when recovery requirements are traceable to design decisions, material gaps have owners, required tests have evidence, operational teams accept the runbooks, and BCM records are updated. This makes resilience part of delivery quality rather than a retrospective audit finding.
Make resilience acceptance part of go-live approval
The go-live decision should clearly state whether continuity requirements are met, partially met or accepted as a temporary exception. Project status should not become “green” simply because functional testing succeeded. The operational owner should know what recovery capability is available on day one, what evidence supports it, and which limitations remain. If the project relies on a future resilience enhancement, record the funded action, owner and due date before the change is accepted.
Apply the same discipline to small changes
Not every change needs a committee review. Use screening rules so routine low-impact changes can pass quickly while changes affecting critical services, recovery architecture, suppliers, data, facilities or emergency arrangements receive deeper review. This keeps the process proportionate and prevents resilience governance from becoming a bottleneck.
Reviewer challenge questions
- Has the change altered a critical service or dependency map?
- Are recovery requirements traceable to design and test evidence?
- Is the alternate mode usable by the people who will actually operate it?
- Have concentration and common-cause failures been considered?
- Can the new service recover without relying on the failed component or location?
- Have support teams, suppliers and business owners accepted the operational model?
Use the answers to decide whether to approve, approve with conditions, defer, or escalate the change. The control is effective when it changes decisions before production, not when it only generates a document after go-live.