RTO and RPO are the two numbers that decide what your recovery plan is allowed to cost, and most organisations set them the way they set a Wi-Fi password — quickly, under mild pressure, and without writing down the reasoning. Someone says “four hours” in a meeting, it goes into a policy document, and three years later an incident discovers that nobody ever priced what four hours would take to build.

This guide fixes that. It explains what each half of RTO and RPO actually measures, where the figures are supposed to come from, and then gives you a business calculator: five inputs you can source from your own finance system, one formula, and an answer that tells you which recovery tier costs your organisation the least once you count both the spend and the risk.

The calculator is deliberately arithmetic you can do in a spreadsheet. There is a worksheet you can copy, a worked example on a sixty-person firm, and indicative UK costs for each tier. If you want the layer definitions underneath all this, our guide to backup vs disaster recovery vs high availability covers what each one protects against before you start putting numbers on it.

One framing note. RTO and RPO are not an infrastructure decision that IT makes and the business ratifies. They are a statement about how much money an hour of outage costs and how much work you can afford to redo, and every technical choice below falls out of those two figures. Teams that start with the technology buy capability they do not need and skip the capability they do.

RTO and RPO explained without the jargon

rto and rpo explained business calculator b three stacked hexagonal plates

Both halves of RTO and RPO are clocks, both start at the moment of failure, and they run in opposite directions. That is the entire concept, and getting RTO and RPO straight in one paragraph each prevents most of the expensive mistakes further down.

RTO is a clock that runs forwards

Recovery Time Objective, the first half of RTO and RPO, is the maximum agreed time between a service failing and that service being usable again. It is a target for duration, and it covers everything: noticing the failure, deciding to invoke the plan, mobilising people, restoring or failing over, validating the result, and telling users they can work again. An RTO of four hours means the business has committed to being operational four hours after the lights go out, not four hours after somebody starts the restore.

RPO is a clock that runs backwards

Recovery Point Objective is the maximum amount of data, expressed as time, that the business is willing to lose. An RPO of one hour means that after recovery you may find yourself back at the state the system was in one hour before the failure, and everything done in that window has to be recreated. RPO is determined almost entirely by how often you copy data, and it is the half of RTO and RPO that people consistently underestimate because it costs work rather than downtime.

Why the two are always quoted as a pair

Either number alone describes half an outage. A service restored in fifteen minutes that has lost a day of transactions is not a successful recovery, and neither is a perfectly current database that takes three days to become reachable. Quoting RTO and RPO together forces the conversation to cover both the wait and the rework, which is why every serious recovery standard treats them as a pair rather than as two independent settings.

The third number nobody sets

There is a third figure that sits above RTO and RPO: maximum tolerable downtime, sometimes called maximum tolerable period of disruption. It is the point at which the damage stops being an inconvenience and becomes structural — a contract breach, a regulatory notification, a cash flow problem, customers who do not come back. Your RTO must be comfortably below it, with margin, because the RTO is a target and the tolerance is a cliff.

DimensionRTORPO
MeasuresTime until the service works againTime worth of data you may lose
Direction from failureForwardsBackwards
Driven byRecovery process and standby capacityCopy frequency and replication method
Cost leverHow warm the second environment isHow often data leaves the primary
Business pain when missedIdle staff, unserved customersRework, reconciliation, lost records
Typical owner of the inputOperations and service leadsFinance and data owners
Cheapest way to improve itRehearse the runbookIncrease snapshot frequency
Most common errorTiming only the restore stepAssuming it equals backup schedule

Where RTO and RPO numbers actually come from

rto and rpo explained business calculator c four rising blank columns

The honest answer is that they come from a business impact analysis, and the reason most organisations skip it is that the phrase sounds like six weeks of consultancy. It does not have to be. For a mid-sized estate you need four figures per service, and three of them already exist in systems you own.

Start with the business service, not the server

Recovery objectives belong to things the business recognises: taking orders, recording billable time, paying people, answering the phone. A server is a component of one of those, and setting RTO and RPO per server produces either an unaffordable estate-wide target or a spreadsheet nobody reads. List the services first, in the language the board uses, and attach the numbers to those.

The hourly cost of downtime is the anchor

For each service, work out what one hour of it being unavailable costs, because this figure anchors both halves of RTO and RPO. Three components cover most of it: the loaded hourly cost of the staff who cannot work, multiplied by how much of their work genuinely stops; the revenue that is lost rather than merely delayed; and any contractual penalty that accrues by the hour. Our breakdown of the cost of downtime and how to calculate the return goes further on sourcing each component defensibly.

The hourly cost of lost data is the other anchor

This one gets forgotten, and it is what makes RPO real. Ask how much work the service records in an hour and what it would cost to recreate it: time entries retyped from memory, orders re-keyed from email, stock counts redone, documents rewritten. Some of it is recoverable at the cost of labour. Some of it is simply gone, and that portion should be priced at its full value rather than at the effort of trying.

Contracts and regulators set floors you cannot negotiate

Before optimising anything, check what you have already promised. Client contracts often contain availability commitments that predate the current infrastructure. Sector regulation may specify notification windows that only make sense if certain systems are back within a defined period. These are floors on your RTO and RPO, and discovering one after an incident is significantly worse than discovering it now.

Failure frequency converts a per-hour cost into an annual one

A cost per hour of outage is not yet a budget. Multiply it by how often you expect a qualifying event: measured from your own incident history where you have it, and from a conservative planning assumption where you do not. Most mid-sized estates land somewhere between one significant disruption every eighteen months and two a year once you include supplier and platform failures. This is the number that turns RTO and RPO from a technical preference into a line a finance director can compare against a premium.

Hourly downtime cost by service — worked example, 60-person professional services firm
Practice management and time recording £2,500
Email and collaboration £1,100
Finance and billing £780
Shared file storage £640
Intranet and internal wiki £90

The spread in that chart is the whole argument against a single estate-wide target. One service is worth twenty-seven times another per hour, and giving both the same RTO and RPO means overspending on one and under-protecting the other, usually at the same time. Tiering RTO and RPO by service is the cheapest saving available in the whole exercise.

The RTO and RPO business calculator

rto and rpo explained business calculator d three upright cylinders row

Here is the calculator itself. Five inputs, three lines of arithmetic, and an answer expressed as total cost of risk — the number that lets you compare a cheap tier with high expected losses against an expensive tier with low ones, on equal terms.

The five inputs

Everything the calculator needs is below. Source each one per service, not per estate, and write down where the figure came from so it can be challenged later. A defensible RTO and RPO pair is one where somebody can trace every number back to a system of record.

Input one: downtime cost per hour, D

D equals the loaded hourly staff cost of the people affected multiplied by the proportion of their work that genuinely stops, plus revenue that is permanently lost per hour rather than delayed, plus any hourly contractual penalty. Take the loaded staff cost from payroll, not from salary — include employer contributions and overheads, which typically add thirty per cent or more. D is the input that does the most work in the RTO and RPO calculation, so it is worth an hour with the finance team.

Input two: data loss cost per hour, L

L equals the volume of records the service creates in an hour multiplied by the cost of recreating each one, plus the value of anything in that hour that cannot be recreated at all. Recreation cost is labour: minutes per record times a loaded hourly rate. The irrecoverable portion is usually small in volume and large in value, so price it separately rather than averaging it away.

Input three: expected failure frequency, F

F is the number of qualifying events you expect per year, expressed as a decimal. Two material incidents in the last three years gives you roughly 0.7. If you have no history, use 0.5 as a starting assumption for a single-site estate with no redundancy, and adjust it down as you add layers. F is the input most worth revisiting annually, and the one that moves RTO and RPO most when a cybersecurity incident lands in the history.

Input four: annual cost of each candidate tier, R

R is the fully loaded annual cost of delivering a given RTO and RPO: licences, storage, standby compute, replication egress, the managed service, and the internal time spent testing it. Testing time is the line most often omitted and it is not small — a properly rehearsed recovery capability costs several days of skilled effort a year before anything fails.

Input five: the tolerance ceiling, T

T is the maximum tolerable downtime for the service, in hours, taken from contracts, regulation, and the point at which cash flow or customer retention breaks. T is not part of the optimisation. It is a veto: any tier whose RTO exceeds T is disqualified no matter how attractive its arithmetic looks, which is why T belongs on the RTO and RPO worksheet rather than in a separate risk register.

The formula

For each candidate tier, expected annual loss is the failure frequency multiplied by the cost of one event, and the cost of one event is the downtime cost times the tier’s RTO plus the data loss cost times the tier’s RPO. Total cost of risk is that expected loss added to the tier’s annual spend. The winning tier is the one with the lowest total, among the tiers that survive the tolerance veto — and that tier’s numbers become your published RTO and RPO.

SymbolInputHow to source itYour value
DDowntime cost per hourPayroll plus revenue analysis plus contract penalties£ ______
LData loss cost per hourRecords per hour times recreation cost, plus irrecoverable value£ ______
FExpected events per yearIncident history, or 0.5 as a starting assumption______
RAnnual cost of the tierLicences, storage, standby compute, egress, testing effort£ ______
TTolerance ceiling, hoursContracts, regulation, cash flow break point______ h
Step 1Cost of one event(D × RTO) + (L × RPO)£ ______
Step 2Expected annual lossF × cost of one event£ ______
Step 3Total cost of riskExpected annual loss + R£ ______
Step 4Veto checkDisqualify any tier where RTO > TPass / fail

Working the RTO and RPO calculator on a real estate

rto and rpo explained business calculator e stepped pyramid four cubes

Abstract formulas persuade nobody. Here is the same calculator run end to end on a plausible sixty-person professional services firm, with every figure shown so you can substitute your own.

The firm and its inputs

Sixty staff, £8.4m annual revenue, one core practice management platform that records billable time and holds client matter files. Loaded staff cost averages £34 an hour. Roughly forty of the sixty cannot do meaningful work when the platform is down. The firm has had two material IT disruptions in three years. Contractual review put the tolerance ceiling at twelve hours, driven by a client obligation to acknowledge instructions same day.

Step one: the hourly numbers

Forty staff at £34 an hour is £1,360, and about seventy per cent of that work genuinely stops, giving £952. Deferred billing that is never recovered is assessed at £1,050 an hour, and there are no hourly penalties. Rounded, D is £2,500 per hour. The platform records around three hundred time entries and document actions an hour, at roughly three minutes each to recreate, so L comes to £900 per hour.

Step two: expected annual loss at each tier

Failure frequency is two events in three years, so F is 0.7. Four candidate tiers were priced, each with its own RTO and RPO pair: backup and restore at 24 hours and 24 hours; pilot light at 4 hours and 1 hour; warm standby at 45 minutes and 5 minutes; and active-active at effectively zero for both. Running the first two lines of the formula against each tier produces the table below.

Step three: total cost of risk

Adding each tier’s annual spend to its expected annual loss gives the comparison the board actually needs. The cheapest tier to buy is not the cheapest tier to own, and neither is the best one. Total cost of risk is the only view in which competing RTO and RPO options can be compared honestly.

TierRTORPOCost of one eventExpected annual lossAnnual spendTotal cost of risk
Backup and restore24 h24 h£81,600£57,120£6,000£63,120
Pilot light4 h1 h£10,900£7,630£19,000£26,630
Warm standby45 min5 min£1,950£1,365£52,000£53,365
Active-activeNear zeroNear zero£0£0£140,000£140,000
Total cost of risk by recovery tier — lower is better
Backup and restore £63,120
Pilot light £26,630
Warm standby £53,365
Active-active £140,000

Step four: the tolerance sanity check

Backup and restore is disqualified before the money is even considered: a 24 hour RTO breaches the twelve hour ceiling. Pilot light wins on total cost of risk and clears the ceiling with eight hours of margin. Warm standby costs twice as much to own for a benefit the firm cannot monetise, and active-active is not a serious candidate at this size.

What the exercise actually changed

The firm had a written RTO of one hour, inherited from a template, and an infrastructure capable of about a day. The calculator did not tighten the target — it relaxed it to four hours, which was both defensible and affordable, and moved the argument from opinion to arithmetic. That is the normal outcome. Honest RTO and RPO numbers usually cost less than aspirational ones.

What each RTO and RPO tier costs to deliver

rto and rpo explained business calculator f upright shield on plinth

The four tiers used above map onto standard recovery patterns, and each one buys a specific RTO and RPO pair rather than a vague sense of safety. Costs below are indicative UK figures for a mid-sized estate and will move with data volume, licensing and how much of the work is in-house, but the ratios between tiers hold up well.

Backup and restore

Data is copied to a second location on a schedule and there is no standby environment at all. Recovery means provisioning somewhere to restore to and then restoring, which is why the RTO is measured in days for anything substantial. It is the right answer for a large share of any estate, and pairing it with an immutable backup strategy is what keeps it credible against ransomware.

Pilot light

Data is replicated continuously into a second region or platform, but the compute is switched off. Recovery means starting instances against data that is already current. RPO drops to minutes because replication is continuous; RTO lands in the low hours because there is still work to do. This is the tier that wins most RTO and RPO calculations at mid-market scale.

Warm standby

A scaled-down but running copy of the environment sits ready and is kept in sync. Failover is a routing change plus a scale-up. RTO falls to tens of minutes and RPO to seconds, at the cost of paying for a second environment that does nothing most of the year. Justified where an hour genuinely costs five figures.

Active-active

Full capacity runs in two locations simultaneously and traffic is served from both. There is no failover event because there is no failure of a single site to recover from. It is the only pattern that delivers near-zero on both halves of RTO and RPO, and it costs roughly what running the estate twice costs, because that is what it is. The lessons from a real cloud region failure are worth reading before assuming it removes all risk.

TierWhat runs at the second siteRealistic RTORealistic RPOIndicative annual UK cost
Backup and restoreNothing, data copies only12 to 48 hours4 to 24 hours£3k to £10k
Pilot lightData replicated, compute off2 to 6 hours5 to 60 minutes£12k to £30k
Warm standbyScaled-down live copy15 to 60 minutesSeconds to 5 minutes£35k to £80k
Active-activeFull production capacityNear zeroNear zero£100k and above

Why the middle tier usually wins

The pattern in that table is not an accident. Expected loss falls fast as RTO drops from a day to a few hours, because most of the damage happens in the first shift. Beyond that, loss falls slowly while cost climbs steeply. The minimum of the combined curve therefore sits in the middle, which is why sensible RTO and RPO targets so often land at a few hours and a few minutes rather than at zero.

Turning RTO and RPO targets into a design that works

A number in a policy document does nothing. Four things turn agreed RTO and RPO targets into a capability that behaves the way the spreadsheet says it will.

Tier the estate, do not average it

Group services into three or four tiers with a defined RTO and RPO each, and assign every service to a tier explicitly. Three tiers is usually enough: critical, important, and deferrable. Anything not assigned defaults to the lowest tier by policy, which forces owners to argue for an upgrade rather than quietly assuming one.

Match the pattern to the number, then check the dependencies

A service inherits the worst RTO of everything it depends on. An application with a four hour target that authenticates against a directory with a twelve hour target has a twelve hour target, whatever the documentation claims. Identity, DNS, network connectivity and shared data platforms belong in the highest tier almost by definition, because everything else waits on them.

Budget for the decision, not just the technology

In most rehearsals the largest single block of the recovery clock is not the restore. It is the gap between something failing and somebody with authority declaring that the plan is now in effect. Pre-agreed invocation criteria and a named decision-maker with a deputy shorten real RTO more cheaply than any infrastructure change, which is a point our guide to business continuity planning for outages develops in detail.

Where the clock goes inside a four-hour RTO — typical rehearsal breakdown, minutes
Detect and escalate 25 min
Decide and invoke 30 min
Restore or fail over 95 min
Validate and reconcile 45 min
Cut users back over 25 min

Write the numbers into the contract

If a supplier delivers any part of the recovery, the agreed RTO and RPO belong in the contract with a defined measurement method and a stated consequence. A vendor availability percentage is not a recovery commitment, and an unqualified promise to “use reasonable endeavours” is not either. Our overview of enterprise disaster recovery solutions covers what a workable clause looks like.

Testing is the only proof your RTO and RPO are real

RTO and RPO that have never been measured are forecasts. The test converts them into facts, and the gap it exposes is almost always in the same places.

Measure the whole clock, not the restore step

Start the stopwatch at simulated failure and stop it when a real user does real work. Teams that time only the technical restore report RTO and RPO figures two to four times better than the business experiences, then get blamed for missing a target they were never actually measuring. Record each phase separately so you know which one to attack.

Test the worst realistic case, not the convenient one

A restore of one small database on a Tuesday afternoon proves very little. Test the scenario your objectives were written for: the primary platform unavailable, the usual engineer on leave, and the credentials in a vault that also lives in the affected environment. Ransomware deserves its own rehearsal because it invalidates recent copies, a dynamic the ransomware defence playbook works through properly.

Record the gap and act on it

Write down measured RTO and RPO against target after every exercise, with the reason for any shortfall and a named owner for the fix. A register with three years of measured RTO and RPO results is worth more to an auditor, a client, and an insurer than any amount of policy documentation, and it makes the next budget conversation straightforward.

Retest after every material change

Recovery capability decays silently. A new integration, a migrated workload, a changed identity provider or a renewed supplier contract can each move your real RTO and RPO without anyone noticing. Retest when the estate changes rather than annually by calendar, and treat a passed test as valid only for the configuration it was run against.

Mistakes that quietly make RTO and RPO meaningless

Every item here has been seen in a real plan whose RTO and RPO looked entirely reasonable on paper.

Setting one target for the whole estate

A single RTO and RPO pair over-protects the intranet and under-protects the billing platform simultaneously, at considerable expense. Tier it.

Confusing backup frequency with RPO

Nightly backups do not give you a 24 hour RPO. They give you an RPO of up to 24 hours plus however long verification and restore staging take before that copy is usable, and that assumes the most recent copy is clean.

Assuming the newest copy is available

Against ransomware the recent copies are frequently the compromised ones, and the usable recovery point may be days older than your stated RPO. Plan for the copy you can trust rather than the copy you have, and set RTO and RPO against that copy.

Ignoring the cost of deciding

If invocation requires a director who is unreachable at weekends, the real RTO includes however long it takes to find them. Authority delegation is free and it buys hours.

Treating a vendor SLA as your RTO

A platform’s uptime commitment describes what the vendor owes you in credits. It is not a recovery time for your business service, and it does not cover data you deleted or corrupted yourself.

Never pricing the alternative

A target chosen without comparing tiers cannot be justified or defended. Running the calculator on two alternatives takes an afternoon and turns RTO and RPO from a preference into a decision with an audit trail.

RTO and RPO FAQ

What is a realistic RTO for a small business?

For core operational systems, four to eight hours is achievable and affordable using a pilot light pattern. Anything under an hour requires continuously running standby capacity, and that cost should be justified by a measured hourly downtime figure rather than by a preference for a round number.

Can RPO ever be zero?

Only with synchronous replication, which requires every write to be committed in two places before the application is told it succeeded. That adds latency, constrains the distance between sites, and does not protect against logical corruption, because a faithful copy of bad data is still bad data.

Is RTO measured from the failure or from the decision?

From the failure. Measuring from the invocation decision is the single most common way organisations report recovery times that the business does not recognise, because it excludes detection, escalation and the wait for someone to take responsibility.

How do RTO and RPO apply to SaaS platforms?

RTO and RPO still apply, but you are inheriting the provider’s numbers rather than setting your own, and the provider’s retention limits define your recovery point for anything you delete or corrupt. That is the standard argument for independent protection of a Microsoft 365 or Google Workspace tenant.

How often should the numbers be reviewed?

Annually as a minimum, and whenever the business changes materially — a new revenue line, an acquisition, a major system migration, or a new client contract with availability terms. The input most worth rechecking is the hourly downtime cost, which tends to rise faster than anyone expects.

Who should sign off RTO and RPO?

The business owner of the service, not IT. IT owns the estimate of what a given RTO and RPO pair costs and whether it is achievable; the business owns the decision about how much loss is acceptable. Recording that signature is what makes the target defensible when it is eventually tested by an actual incident.

References