RTO and RPO are the two numbers that decide what your recovery plan is allowed to cost, and most organisations set them the way they set a Wi-Fi password — quickly, under mild pressure, and without writing down the reasoning. Someone says “four hours” in a meeting, it goes into a policy document, and three years later an incident discovers that nobody ever priced what four hours would take to build.
This guide fixes that. It explains what each half of RTO and RPO actually measures, where the figures are supposed to come from, and then gives you a business calculator: five inputs you can source from your own finance system, one formula, and an answer that tells you which recovery tier costs your organisation the least once you count both the spend and the risk.
The calculator is deliberately arithmetic you can do in a spreadsheet. There is a worksheet you can copy, a worked example on a sixty-person firm, and indicative UK costs for each tier. If you want the layer definitions underneath all this, our guide to backup vs disaster recovery vs high availability covers what each one protects against before you start putting numbers on it.
One framing note. RTO and RPO are not an infrastructure decision that IT makes and the business ratifies. They are a statement about how much money an hour of outage costs and how much work you can afford to redo, and every technical choice below falls out of those two figures. Teams that start with the technology buy capability they do not need and skip the capability they do.
Table of contents
- RTO and RPO explained without the jargon
- Where RTO and RPO numbers actually come from
- The RTO and RPO business calculator
- Working the RTO and RPO calculator on a real estate
- What each RTO and RPO tier costs to deliver
- Turning RTO and RPO targets into a design that works
- Testing is the only proof your RTO and RPO are real
- Mistakes that quietly make RTO and RPO meaningless
- RTO and RPO FAQ
- References
RTO and RPO explained without the jargon
Both halves of RTO and RPO are clocks, both start at the moment of failure, and they run in opposite directions. That is the entire concept, and getting RTO and RPO straight in one paragraph each prevents most of the expensive mistakes further down.
RTO is a clock that runs forwards
Recovery Time Objective, the first half of RTO and RPO, is the maximum agreed time between a service failing and that service being usable again. It is a target for duration, and it covers everything: noticing the failure, deciding to invoke the plan, mobilising people, restoring or failing over, validating the result, and telling users they can work again. An RTO of four hours means the business has committed to being operational four hours after the lights go out, not four hours after somebody starts the restore.
RPO is a clock that runs backwards
Recovery Point Objective is the maximum amount of data, expressed as time, that the business is willing to lose. An RPO of one hour means that after recovery you may find yourself back at the state the system was in one hour before the failure, and everything done in that window has to be recreated. RPO is determined almost entirely by how often you copy data, and it is the half of RTO and RPO that people consistently underestimate because it costs work rather than downtime.
Why the two are always quoted as a pair
Either number alone describes half an outage. A service restored in fifteen minutes that has lost a day of transactions is not a successful recovery, and neither is a perfectly current database that takes three days to become reachable. Quoting RTO and RPO together forces the conversation to cover both the wait and the rework, which is why every serious recovery standard treats them as a pair rather than as two independent settings.
The third number nobody sets
There is a third figure that sits above RTO and RPO: maximum tolerable downtime, sometimes called maximum tolerable period of disruption. It is the point at which the damage stops being an inconvenience and becomes structural — a contract breach, a regulatory notification, a cash flow problem, customers who do not come back. Your RTO must be comfortably below it, with margin, because the RTO is a target and the tolerance is a cliff.
| Dimension | RTO | RPO |
|---|---|---|
| Measures | Time until the service works again | Time worth of data you may lose |
| Direction from failure | Forwards | Backwards |
| Driven by | Recovery process and standby capacity | Copy frequency and replication method |
| Cost lever | How warm the second environment is | How often data leaves the primary |
| Business pain when missed | Idle staff, unserved customers | Rework, reconciliation, lost records |
| Typical owner of the input | Operations and service leads | Finance and data owners |
| Cheapest way to improve it | Rehearse the runbook | Increase snapshot frequency |
| Most common error | Timing only the restore step | Assuming it equals backup schedule |
Where RTO and RPO numbers actually come from
The honest answer is that they come from a business impact analysis, and the reason most organisations skip it is that the phrase sounds like six weeks of consultancy. It does not have to be. For a mid-sized estate you need four figures per service, and three of them already exist in systems you own.
Start with the business service, not the server
Recovery objectives belong to things the business recognises: taking orders, recording billable time, paying people, answering the phone. A server is a component of one of those, and setting RTO and RPO per server produces either an unaffordable estate-wide target or a spreadsheet nobody reads. List the services first, in the language the board uses, and attach the numbers to those.
The hourly cost of downtime is the anchor
For each service, work out what one hour of it being unavailable costs, because this figure anchors both halves of RTO and RPO. Three components cover most of it: the loaded hourly cost of the staff who cannot work, multiplied by how much of their work genuinely stops; the revenue that is lost rather than merely delayed; and any contractual penalty that accrues by the hour. Our breakdown of the cost of downtime and how to calculate the return goes further on sourcing each component defensibly.
The hourly cost of lost data is the other anchor
This one gets forgotten, and it is what makes RPO real. Ask how much work the service records in an hour and what it would cost to recreate it: time entries retyped from memory, orders re-keyed from email, stock counts redone, documents rewritten. Some of it is recoverable at the cost of labour. Some of it is simply gone, and that portion should be priced at its full value rather than at the effort of trying.
Contracts and regulators set floors you cannot negotiate
Before optimising anything, check what you have already promised. Client contracts often contain availability commitments that predate the current infrastructure. Sector regulation may specify notification windows that only make sense if certain systems are back within a defined period. These are floors on your RTO and RPO, and discovering one after an incident is significantly worse than discovering it now.
Failure frequency converts a per-hour cost into an annual one
A cost per hour of outage is not yet a budget. Multiply it by how often you expect a qualifying event: measured from your own incident history where you have it, and from a conservative planning assumption where you do not. Most mid-sized estates land somewhere between one significant disruption every eighteen months and two a year once you include supplier and platform failures. This is the number that turns RTO and RPO from a technical preference into a line a finance director can compare against a premium.
The spread in that chart is the whole argument against a single estate-wide target. One service is worth twenty-seven times another per hour, and giving both the same RTO and RPO means overspending on one and under-protecting the other, usually at the same time. Tiering RTO and RPO by service is the cheapest saving available in the whole exercise.
The RTO and RPO business calculator
Here is the calculator itself. Five inputs, three lines of arithmetic, and an answer expressed as total cost of risk — the number that lets you compare a cheap tier with high expected losses against an expensive tier with low ones, on equal terms.
The five inputs
Everything the calculator needs is below. Source each one per service, not per estate, and write down where the figure came from so it can be challenged later. A defensible RTO and RPO pair is one where somebody can trace every number back to a system of record.
Input one: downtime cost per hour, D
D equals the loaded hourly staff cost of the people affected multiplied by the proportion of their work that genuinely stops, plus revenue that is permanently lost per hour rather than delayed, plus any hourly contractual penalty. Take the loaded staff cost from payroll, not from salary — include employer contributions and overheads, which typically add thirty per cent or more. D is the input that does the most work in the RTO and RPO calculation, so it is worth an hour with the finance team.
Input two: data loss cost per hour, L
L equals the volume of records the service creates in an hour multiplied by the cost of recreating each one, plus the value of anything in that hour that cannot be recreated at all. Recreation cost is labour: minutes per record times a loaded hourly rate. The irrecoverable portion is usually small in volume and large in value, so price it separately rather than averaging it away.
Input three: expected failure frequency, F
F is the number of qualifying events you expect per year, expressed as a decimal. Two material incidents in the last three years gives you roughly 0.7. If you have no history, use 0.5 as a starting assumption for a single-site estate with no redundancy, and adjust it down as you add layers. F is the input most worth revisiting annually, and the one that moves RTO and RPO most when a cybersecurity incident lands in the history.
Input four: annual cost of each candidate tier, R
R is the fully loaded annual cost of delivering a given RTO and RPO: licences, storage, standby compute, replication egress, the managed service, and the internal time spent testing it. Testing time is the line most often omitted and it is not small — a properly rehearsed recovery capability costs several days of skilled effort a year before anything fails.
Input five: the tolerance ceiling, T
T is the maximum tolerable downtime for the service, in hours, taken from contracts, regulation, and the point at which cash flow or customer retention breaks. T is not part of the optimisation. It is a veto: any tier whose RTO exceeds T is disqualified no matter how attractive its arithmetic looks, which is why T belongs on the RTO and RPO worksheet rather than in a separate risk register.
The formula
For each candidate tier, expected annual loss is the failure frequency multiplied by the cost of one event, and the cost of one event is the downtime cost times the tier’s RTO plus the data loss cost times the tier’s RPO. Total cost of risk is that expected loss added to the tier’s annual spend. The winning tier is the one with the lowest total, among the tiers that survive the tolerance veto — and that tier’s numbers become your published RTO and RPO.
| Symbol | Input | How to source it | Your value |
|---|---|---|---|
| D | Downtime cost per hour | Payroll plus revenue analysis plus contract penalties | £ ______ |
| L | Data loss cost per hour | Records per hour times recreation cost, plus irrecoverable value | £ ______ |
| F | Expected events per year | Incident history, or 0.5 as a starting assumption | ______ |
| R | Annual cost of the tier | Licences, storage, standby compute, egress, testing effort | £ ______ |
| T | Tolerance ceiling, hours | Contracts, regulation, cash flow break point | ______ h |
| Step 1 | Cost of one event | (D × RTO) + (L × RPO) | £ ______ |
| Step 2 | Expected annual loss | F × cost of one event | £ ______ |
| Step 3 | Total cost of risk | Expected annual loss + R | £ ______ |
| Step 4 | Veto check | Disqualify any tier where RTO > T | Pass / fail |
Working the RTO and RPO calculator on a real estate
Abstract formulas persuade nobody. Here is the same calculator run end to end on a plausible sixty-person professional services firm, with every figure shown so you can substitute your own.
The firm and its inputs
Sixty staff, £8.4m annual revenue, one core practice management platform that records billable time and holds client matter files. Loaded staff cost averages £34 an hour. Roughly forty of the sixty cannot do meaningful work when the platform is down. The firm has had two material IT disruptions in three years. Contractual review put the tolerance ceiling at twelve hours, driven by a client obligation to acknowledge instructions same day.
Step one: the hourly numbers
Forty staff at £34 an hour is £1,360, and about seventy per cent of that work genuinely stops, giving £952. Deferred billing that is never recovered is assessed at £1,050 an hour, and there are no hourly penalties. Rounded, D is £2,500 per hour. The platform records around three hundred time entries and document actions an hour, at roughly three minutes each to recreate, so L comes to £900 per hour.
Step two: expected annual loss at each tier
Failure frequency is two events in three years, so F is 0.7. Four candidate tiers were priced, each with its own RTO and RPO pair: backup and restore at 24 hours and 24 hours; pilot light at 4 hours and 1 hour; warm standby at 45 minutes and 5 minutes; and active-active at effectively zero for both. Running the first two lines of the formula against each tier produces the table below.
Step three: total cost of risk
Adding each tier’s annual spend to its expected annual loss gives the comparison the board actually needs. The cheapest tier to buy is not the cheapest tier to own, and neither is the best one. Total cost of risk is the only view in which competing RTO and RPO options can be compared honestly.
| Tier | RTO | RPO | Cost of one event | Expected annual loss | Annual spend | Total cost of risk |
|---|---|---|---|---|---|---|
| Backup and restore | 24 h | 24 h | £81,600 | £57,120 | £6,000 | £63,120 |
| Pilot light | 4 h | 1 h | £10,900 | £7,630 | £19,000 | £26,630 |
| Warm standby | 45 min | 5 min | £1,950 | £1,365 | £52,000 | £53,365 |
| Active-active | Near zero | Near zero | £0 | £0 | £140,000 | £140,000 |
Step four: the tolerance sanity check
Backup and restore is disqualified before the money is even considered: a 24 hour RTO breaches the twelve hour ceiling. Pilot light wins on total cost of risk and clears the ceiling with eight hours of margin. Warm standby costs twice as much to own for a benefit the firm cannot monetise, and active-active is not a serious candidate at this size.
What the exercise actually changed
The firm had a written RTO of one hour, inherited from a template, and an infrastructure capable of about a day. The calculator did not tighten the target — it relaxed it to four hours, which was both defensible and affordable, and moved the argument from opinion to arithmetic. That is the normal outcome. Honest RTO and RPO numbers usually cost less than aspirational ones.
What each RTO and RPO tier costs to deliver
The four tiers used above map onto standard recovery patterns, and each one buys a specific RTO and RPO pair rather than a vague sense of safety. Costs below are indicative UK figures for a mid-sized estate and will move with data volume, licensing and how much of the work is in-house, but the ratios between tiers hold up well.
Backup and restore
Data is copied to a second location on a schedule and there is no standby environment at all. Recovery means provisioning somewhere to restore to and then restoring, which is why the RTO is measured in days for anything substantial. It is the right answer for a large share of any estate, and pairing it with an immutable backup strategy is what keeps it credible against ransomware.
Pilot light
Data is replicated continuously into a second region or platform, but the compute is switched off. Recovery means starting instances against data that is already current. RPO drops to minutes because replication is continuous; RTO lands in the low hours because there is still work to do. This is the tier that wins most RTO and RPO calculations at mid-market scale.
Warm standby
A scaled-down but running copy of the environment sits ready and is kept in sync. Failover is a routing change plus a scale-up. RTO falls to tens of minutes and RPO to seconds, at the cost of paying for a second environment that does nothing most of the year. Justified where an hour genuinely costs five figures.
Active-active
Full capacity runs in two locations simultaneously and traffic is served from both. There is no failover event because there is no failure of a single site to recover from. It is the only pattern that delivers near-zero on both halves of RTO and RPO, and it costs roughly what running the estate twice costs, because that is what it is. The lessons from a real cloud region failure are worth reading before assuming it removes all risk.
| Tier | What runs at the second site | Realistic RTO | Realistic RPO | Indicative annual UK cost |
|---|---|---|---|---|
| Backup and restore | Nothing, data copies only | 12 to 48 hours | 4 to 24 hours | £3k to £10k |
| Pilot light | Data replicated, compute off | 2 to 6 hours | 5 to 60 minutes | £12k to £30k |
| Warm standby | Scaled-down live copy | 15 to 60 minutes | Seconds to 5 minutes | £35k to £80k |
| Active-active | Full production capacity | Near zero | Near zero | £100k and above |
Why the middle tier usually wins
The pattern in that table is not an accident. Expected loss falls fast as RTO drops from a day to a few hours, because most of the damage happens in the first shift. Beyond that, loss falls slowly while cost climbs steeply. The minimum of the combined curve therefore sits in the middle, which is why sensible RTO and RPO targets so often land at a few hours and a few minutes rather than at zero.
Turning RTO and RPO targets into a design that works
A number in a policy document does nothing. Four things turn agreed RTO and RPO targets into a capability that behaves the way the spreadsheet says it will.
Tier the estate, do not average it
Group services into three or four tiers with a defined RTO and RPO each, and assign every service to a tier explicitly. Three tiers is usually enough: critical, important, and deferrable. Anything not assigned defaults to the lowest tier by policy, which forces owners to argue for an upgrade rather than quietly assuming one.
Match the pattern to the number, then check the dependencies
A service inherits the worst RTO of everything it depends on. An application with a four hour target that authenticates against a directory with a twelve hour target has a twelve hour target, whatever the documentation claims. Identity, DNS, network connectivity and shared data platforms belong in the highest tier almost by definition, because everything else waits on them.
Budget for the decision, not just the technology
In most rehearsals the largest single block of the recovery clock is not the restore. It is the gap between something failing and somebody with authority declaring that the plan is now in effect. Pre-agreed invocation criteria and a named decision-maker with a deputy shorten real RTO more cheaply than any infrastructure change, which is a point our guide to business continuity planning for outages develops in detail.
Write the numbers into the contract
If a supplier delivers any part of the recovery, the agreed RTO and RPO belong in the contract with a defined measurement method and a stated consequence. A vendor availability percentage is not a recovery commitment, and an unqualified promise to “use reasonable endeavours” is not either. Our overview of enterprise disaster recovery solutions covers what a workable clause looks like.
Testing is the only proof your RTO and RPO are real
RTO and RPO that have never been measured are forecasts. The test converts them into facts, and the gap it exposes is almost always in the same places.
Measure the whole clock, not the restore step
Start the stopwatch at simulated failure and stop it when a real user does real work. Teams that time only the technical restore report RTO and RPO figures two to four times better than the business experiences, then get blamed for missing a target they were never actually measuring. Record each phase separately so you know which one to attack.
Test the worst realistic case, not the convenient one
A restore of one small database on a Tuesday afternoon proves very little. Test the scenario your objectives were written for: the primary platform unavailable, the usual engineer on leave, and the credentials in a vault that also lives in the affected environment. Ransomware deserves its own rehearsal because it invalidates recent copies, a dynamic the ransomware defence playbook works through properly.
Record the gap and act on it
Write down measured RTO and RPO against target after every exercise, with the reason for any shortfall and a named owner for the fix. A register with three years of measured RTO and RPO results is worth more to an auditor, a client, and an insurer than any amount of policy documentation, and it makes the next budget conversation straightforward.
Retest after every material change
Recovery capability decays silently. A new integration, a migrated workload, a changed identity provider or a renewed supplier contract can each move your real RTO and RPO without anyone noticing. Retest when the estate changes rather than annually by calendar, and treat a passed test as valid only for the configuration it was run against.
Mistakes that quietly make RTO and RPO meaningless
Every item here has been seen in a real plan whose RTO and RPO looked entirely reasonable on paper.
Setting one target for the whole estate
A single RTO and RPO pair over-protects the intranet and under-protects the billing platform simultaneously, at considerable expense. Tier it.
Confusing backup frequency with RPO
Nightly backups do not give you a 24 hour RPO. They give you an RPO of up to 24 hours plus however long verification and restore staging take before that copy is usable, and that assumes the most recent copy is clean.
Assuming the newest copy is available
Against ransomware the recent copies are frequently the compromised ones, and the usable recovery point may be days older than your stated RPO. Plan for the copy you can trust rather than the copy you have, and set RTO and RPO against that copy.
Ignoring the cost of deciding
If invocation requires a director who is unreachable at weekends, the real RTO includes however long it takes to find them. Authority delegation is free and it buys hours.
Treating a vendor SLA as your RTO
A platform’s uptime commitment describes what the vendor owes you in credits. It is not a recovery time for your business service, and it does not cover data you deleted or corrupted yourself.
Never pricing the alternative
A target chosen without comparing tiers cannot be justified or defended. Running the calculator on two alternatives takes an afternoon and turns RTO and RPO from a preference into a decision with an audit trail.
RTO and RPO FAQ
What is a realistic RTO for a small business?
For core operational systems, four to eight hours is achievable and affordable using a pilot light pattern. Anything under an hour requires continuously running standby capacity, and that cost should be justified by a measured hourly downtime figure rather than by a preference for a round number.
Can RPO ever be zero?
Only with synchronous replication, which requires every write to be committed in two places before the application is told it succeeded. That adds latency, constrains the distance between sites, and does not protect against logical corruption, because a faithful copy of bad data is still bad data.
Is RTO measured from the failure or from the decision?
From the failure. Measuring from the invocation decision is the single most common way organisations report recovery times that the business does not recognise, because it excludes detection, escalation and the wait for someone to take responsibility.
How do RTO and RPO apply to SaaS platforms?
RTO and RPO still apply, but you are inheriting the provider’s numbers rather than setting your own, and the provider’s retention limits define your recovery point for anything you delete or corrupt. That is the standard argument for independent protection of a Microsoft 365 or Google Workspace tenant.
How often should the numbers be reviewed?
Annually as a minimum, and whenever the business changes materially — a new revenue line, an acquisition, a major system migration, or a new client contract with availability terms. The input most worth rechecking is the hourly downtime cost, which tends to rise faster than anyone expects.
Who should sign off RTO and RPO?
The business owner of the service, not IT. IT owns the estimate of what a given RTO and RPO pair costs and whether it is achievable; the business owns the decision about how much loss is acceptable. Recording that signature is what makes the target defensible when it is eventually tested by an actual incident.
References
NIST SP 800-34 Rev. 1, Contingency Planning Guide for Federal Information Systems
NIST SP 1800-11, Data Integrity: Recovering from Ransomware and Other Destructive Events
ISO 22301:2019, Security and resilience — Business continuity management systems
Business continuity and disaster recovery planning — NCSC
Mitigating malware and ransomware attacks — NCSC
Disaster recovery options in the cloud — AWS Whitepapers
AWS Well-Architected Framework, Reliability Pillar
Recommendations for designing a disaster recovery strategy — Azure Well-Architected Framework
About Azure Site Recovery — Microsoft Learn
Disaster recovery planning guide — Google Cloud Architecture Center
Availability table — Google Site Reliability Engineering