As of 11 October 2026. Figures in the worked examples are illustrative. Replace them with measurements from your own restore tests.
Why most recovery targets are wishes
Ask a team for its recovery time objective and the answer is usually a round number agreed in a meeting: “four hours”. Ask when that number was last proven by a timed restore and the room goes quiet. A recovery target that has never been measured is a wish. This guide shows how to set RTO and RPO from business impact, translate them into backup and restore engineering, and then prove them with tests you can repeat.
The three numbers, defined
NIST SP 800-34 Rev. 1, the US federal contingency planning guide, defines the terms most organisations use:
| Term | NIST SP 800-34 Rev. 1 definition (summarised) | Who owns it |
|---|---|---|
| Maximum Tolerable Downtime (MTD) | How long a mission or business process can be disrupted without causing significant harm to the organisation | The business process owner |
| Recovery Time Objective (RTO) | How long a system’s components can be in the recovery phase before the outage harms the mission or business processes | The system owner, within the MTD |
| Recovery Point Objective (RPO) | The point in time to which data must be recovered after an outage | The data owner |
Two rules follow from the definitions. First, the RTO of every system a process depends on has to fit inside that process’s MTD, with time left for the people side of recovery. Second, RTO and RPO are business decisions expressed in engineering terms; IT measures whether they are achievable, but it should not invent them.
RPO is not your backup schedule
A nightly backup does not give you a 24-hour RPO. The data you can actually lose is wider than the interval between jobs:
worst-case data loss ≈ backup interval
+ backup job duration (data written during the job is not in it)
+ copy / replication lag to the recovery location
+ one more interval if the last job failed unnoticed
Illustrative example: nightly backups at 23:00 that take 3 hours, copied off-site with a 2-hour lag, give a realistic worst case of about 29 hours, not 24. If a failed job goes unnoticed for a day, it is about 53 hours. That is why backup job alerting is part of meeting an RPO, not a separate task.
For ransomware, add one more factor: the newest restore point may already contain the attacker’s changes. Finding a clean point can push recovery back by days, which is a reason to keep immutable copies across a longer window, and to rehearse choosing a clean point.
RTO is more than restore speed
Vendors quote restore throughput. Your RTO is the whole recovery phase:
| Phase | What happens | Typical evidence to time |
|---|---|---|
| Detect | Monitoring or a user notices the outage | Alert timestamp vs. first ticket |
| Decide | Someone with authority declares recovery and picks the restore point | Incident log |
| Provision | Target hardware, hypervisor or cloud capacity is ready | Change log |
| Restore | Data is copied or mounted from backup | Backup software job report |
| Validate | Application owners confirm the system works and data is consistent | Sign-off record |
| Reconnect | Users, integrations and DNS point to the recovered system | Change log, first successful transaction |
Illustrative example: restoring 4 TB at a sustained 400 MB/s takes about 10,000 seconds, roughly 2 hours 45 minutes. Add an hour each to detect, decide and validate, and the honest RTO is closer to 6 hours than the 3 hours the restore alone suggests. Instant-recovery features that run a VM directly from backup storage can shorten the restore phase, but performance while running from backup storage needs testing too.
Set targets by tier, not by system
Few organisations can afford short targets for everything. Group systems into tiers, agree targets per tier with the business, and choose the protection technique that can meet each one. The values below are an example starting point, not a standard.
| Example tier | Example systems | Example RTO / RPO | Techniques that can meet it |
|---|---|---|---|
| Tier 1: critical | ERP, core banking or billing, identity | 1–4 h / 15 min–1 h | Replication to a second site, frequent snapshots or log shipping, instant recovery |
| Tier 2: important | Email, file services, CRM | 8–24 h / 4–12 h | Several backups a day, local plus off-site copies |
| Tier 3: deferrable | Archives, test systems, reporting | 2–5 days / 24 h | Daily backup, off-site and immutable copy |
Put a price on the gap
The XOOPIE RTO & RPO calculator gives a first-order cost of an outage from five inputs. With its default inputs (50 employees, ₹1,200 per employee-hour, ₹10 crore annual revenue, a 24-hour backup interval and a 12-hour recovery time), it estimates:
- Revenue at risk of ₹50,000 an hour (annual revenue spread over 250 eight-hour working days) plus ₹60,000 an hour of idle staff time: ₹1,10,000 per hour of downtime.
- Downtime loss over 12 hours: ₹13.2 lakh.
- Cost of recreating 24 hours of lost work (75% of staff cost per lost hour): ₹10.8 lakh.
- Total: about ₹24 lakh per incident.
Moving the same organisation to a 4-hour backup interval and a 4-hour recovery time brings the estimate down to about ₹6.2 lakh. The difference is the budget envelope for better protection. The model is deliberately simple: it ignores penalties, reputational cost and seasonal peaks, so treat it as a starting point for the business conversation, not a valuation.
Prove it: a repeatable restore test
- Pick one system per tier each quarter (our recommendation), rotating through the estate during the year.
- Restore into an isolated network so the test cannot affect production, and time every phase in the table above.
- Record the restore point you used and calculate the data you would have lost. Compare it with the RPO.
- Have the application owner validate with a scripted check, for example logging in and running a known report.
- Write down the gaps: missing runbook steps, passwords no one had, licences that would not activate in the recovery site.
- Retest after fixing, then update the tier targets or the protection technique if the measurement keeps missing.
Keep the evidence. Auditors, cyber insurers and your own management will ask for proof that recovery works, and a dated test record with measured times is that proof. For the practicalities of restore testing, see why backup testing matters and ransomware recovery planning.
Sources
Primary sources used for this article (checked 11 October 2026):
- NIST SP 800-34 Rev. 1: Contingency Planning Guide for Federal Information Systems
- NIST glossary: Recovery Time Objective
- NIST glossary: Recovery Point Objective
- NIST glossary: Maximum Tolerable Downtime