
Backup software reports on itself. A successful job means data was read, written and, where configured, verified. It says nothing about whether a working service can be rebuilt from that copy, in what order, by whom, and inside the time you can survive without it.
A backup is a stored artefact. A recovery is an operation carried out by people under pressure. The first is bought. The second is rehearsed.
Recovery objectives are decisions, not settings
Recovery objectives get typed into a console as though they were product settings. They are business decisions, owned by whoever is accountable for the service, not by the platform team.
A recovery point objective states how much work you are prepared to redo. Set it at four hours for a finance system and you are saying a morning of postings may be re-keyed. Someone should agree to that in those words. A four-hour objective served by a nightly backup is a wish.
A recovery time objective states how long you can operate without the service. Replication interval is not recovery time. Measure from the moment the service stopped, not from the moment someone decided to act: detection, the restore, the client reconfiguration, the reconciliation and the verification all sit inside the window.
Set objectives per system. One policy for everything guarantees that something critical is under-protected and something trivial is expensive.
What a drill exposes that a document cannot
Dependency order breaks first. Directory, DNS, certificate and time services are needed before anything authenticates, and the identity platform is often waiting in the same queue. Kerberos rejects tickets when clocks differ by more than a few minutes, so a domain controller restored with the wrong time authenticates nobody. The virtualisation management platform is sometimes running on the cluster it is meant to bring back.
Credentials break second. The break-glass account sits in a password manager behind single sign-on that depends on the directory that is down. The passphrase for the backup repository lives in one engineer's records, not the organisation's.
Authority breaks third. In a real incident the pull is to keep repairing for another hour, then another. A plan that does not name who can declare a disaster, and the threshold for doing so, loses those hours before recovery starts.
The runbook fails before the platform does
Restore documentation written by the installing supplier describes the product. It assumes the original addressing, names hosts that have since been rebuilt, and does not know which of your services must come back first.
A useful test is scheduled, timed and run by the staff who would run it for real. It restores into an isolated network, so a recovered domain controller or DHCP server cannot collide with the live one. It covers a single file, a full application and a complete site, because those are three different operations. Whoever runs it follows the runbook literally, and every question they have to ask is a defect in the document.
When you assess a supplier, ask for the date of the last full failover test, the recovery time it measured and what changed afterwards: the runbook, the objective, or both. If those do not exist, what you have been sold is backup.
Continue reading
Next step
Bring the complete environment into one conversation.
Tell us what you are planning, replacing, integrating or trying to stabilise. We will help define the right next step.

