Metrics

BCM KPI Framework: Define Decision-Grade Continuity Measures

A measurement-design framework for defining BCM KPIs with auditable formulas, data ownership, thresholds and management actions.

BCM metrics are useful when their definitions are precise enough for another reviewer to reproduce the result and when a threshold causes a named management action. This framework focuses on designing the measures themselves, not on dashboard layout or executive reporting.

Start with the decision, not the percentage

For every proposed KPI, state the management question first. Examples include: Are critical services demonstrably recoverable within approved tolerances? Are material continuity gaps being closed before risk acceptance expires? Are critical suppliers providing current recovery evidence? If a metric does not change a decision, remove it from the core KPI set.

Metric specification

FieldWhat to define
PurposeThe decision or control the metric supports.
FormulaNumerator, denominator, exclusions and treatment of missing data.
SourceAuthoritative record and evidence location.
OwnerPerson accountable for data quality and interpretation.
FrequencyRefresh and independent validation cycle.
ThresholdGreen, amber and red conditions tied to approved tolerance.
ResponseAction, escalation or risk decision required after breach.

Separate four types of measure

  • Coverage: whether required BCM scope is represented.
  • Currency: whether analysis, plans and evidence remain valid after change.
  • Capability: what exercises or recovery tests actually demonstrated.
  • Exposure: unresolved gaps, concentration risks and overdue corrective actions.

Do not allow high coverage to compensate for failed capability. A portfolio with 100% plans but only 60% of critical services tested end-to-end should not appear healthy.

Worked KPI definition

Critical services proven within RTO: numerator = critical services completing an end-to-end recovery test within the approved review period with achieved recovery time at or below approved RTO; denominator = all currently approved critical services. Desktop walkthroughs that did not execute critical dependencies are excluded from the numerator, not the denominator. Amber triggers a remediation plan; red requires executive escalation or documented risk acceptance.

Data-quality controls

  • Retain calculation logic and source evidence for each reporting period.
  • Show missing records instead of silently removing them from the denominator.
  • Distinguish self-reported readiness from independently validated evidence.
  • Record material scope changes that alter the denominator.
  • Prevent manually overridden status without an approver and reason.

Acceptance test for a BCM KPI

A competent reviewer should be able to reproduce the value from retained evidence, understand why the threshold exists, identify who must act after a breach and trace the resulting decision. Use the BCM Program Health Dashboard when these defined measures need to be assembled into an executive view, and the BCM KPI/KRI guide when deciding whether an indicator represents performance or changing risk.

Start with the decision each metric supports

A continuity KPI should answer a management question. Examples include whether critical recovery capability is improving, whether overdue corrective actions are concentrating risk, or whether tested recovery times remain inside approved targets. If a metric does not lead to a decision, escalation or investigation, it is probably reporting activity rather than resilience.

Define numerator, denominator and cut-off

Document the exact calculation, population, exclusions, reporting cut-off and data owner. “Plans tested: 90%” is weak unless the denominator states which plans were due, how partial tests are treated and whether expired evidence counts. Freeze the reporting population at a defined cut-off so historical results can be reproduced instead of changing when records are later edited.

Weight criticality transparently

Do not let many low-criticality successes hide one critical failure. Report critical services separately or use a documented weighting method that management can understand. Pair percentages with counts and named exceptions. A green average should never conceal an untested service whose MTPD is measured in hours.

Combine leading and lagging indicators

Leading indicators can include exercise coverage, recovery evidence age, overdue dependency reviews, alternate-site readiness and corrective-action closure. Lagging indicators can include actual recovery performance, incident workaround failure, missed RTO/RPO and disruption losses. Use both: leading indicators show deteriorating preparedness while lagging indicators test whether preparedness worked when needed.

Set thresholds and escalation rules

For each KPI define green, amber and red criteria, trend tolerance and the action expected when a threshold is breached. Repeated amber performance may deserve escalation even when no single month is red. Record who can accept an exception and for how long. Dashboards should link from the indicator to underlying evidence, failed services and corrective actions so management can investigate rather than debate the number.

Protect metric integrity

Assign an accountable data owner and an independent reviewer for important measures. Keep source timestamps, evidence links and calculation versions. Review metrics after organizational changes so denominators do not silently become stale. Periodically retire metrics that no longer drive decisions and add measures for newly material risks. The objective is a small set of trusted, decision-grade indicators rather than a large dashboard of easy-to-count activity.

Example decision-grade measures

For recovery capability, report the percentage of critical services with a successful end-to-end exercise within the required cycle, alongside the count of critical services that failed RTO or RPO. For corrective actions, show overdue high-severity actions as a count and percentage, the oldest overdue age, and whether any relate to services with short MTPD. For dependency assurance, report critical third parties with current evidence and separately flag those whose evidence is older than the defined review period.

Trend interpretation matters. A move from 92% to 95% exercise coverage may look positive while the remaining 5% contains the organization’s two most critical services. The dashboard should therefore present the exception population and business consequence beside the headline measure. Management review should record the resulting decision: accept temporarily, fund remediation, change a target, or require an accelerated test.