RTO vs RPO: Setting Recovery Targets for Business-Critical Software

Content authorBy Lincoln WoolseyPublished onReading time14 min read
A modern office space with an IT consultant discussing disaster recovery plans, featuring a laptop and collaborative professionals in the background.

Every resilience conversation eventually needs to answer two questions: how long can this system be down, and how much data can you afford to lose? Here's how to turn those questions into numbers the board will actually sign off on.

Why recovery targets get set too late

Most disaster recovery planning conversations start after the invoice arrives. The order management system goes down on a Tuesday and sales stalls for six hours before someone asks the question that should have been answered eighteen months earlier: how long were we supposed to be able to survive this? That is when recovery time objective vs recovery point objective stops being an IT acronym and becomes a commercial argument about who signed off on the exposure.

The numbers make the case. Information Technology Intelligence Consulting polled over 1,000 firms and found a single hour of downtime costs more than $300,000 for over 90% of mid-size and large enterprises. Uptime Institute's survey put one in five recent outages above $1 million. You cannot buy business continuity sensibly until you know what an hour costs you.

Recovery time objective vs recovery point objective

Recovery Time Objective (RTO) is the maximum period a service can be unavailable before the disruption becomes unacceptable to the business. Recovery Point Objective (RPO) is the maximum amount of data, expressed as a period of time, that you can afford to lose permanently. Take one outage timeline: your platform fails at 09:00 and is fully usable again at 13:00. That four-hour gap is measured against your RTO.

Now look backwards from the failure. If your last usable backup completed at 06:00, then three hours of orders and customer records are gone and have to be reconstructed by hand or written off. That three-hour gap is measured against your RPO. Recovery time objective vs recovery point objective refer to the same incident, but they describe different losses.

The distinction that matters commercially is this: recovery time objective vs recovery point objective are business thresholds. A hosting provider quoting 99.99% availability is describing what it expects its infrastructure to do. Your RTO describes what your finance director will tolerate before invoicing stops. Those two things are set by different people for different reasons, and confusing them is how organisations end up paying for resilience they don't need while lacking the resilience they do.

The Financial Conduct Authority formalised exactly this thinking for regulated firms. Under Policy Statement PS21/3, firms must set impact tolerances for each important business service, defined as the maximum tolerable disruption measured in time. The regulator deliberately requires that tolerance to be set without reference to existing recovery capability. You decide what the business can bear first, then find out whether you can afford it.

Compare business consequences

Missing your RTO and missing your RPO fail in different directions, which is why a single "resilience budget" for recovery time objective vs recovery point objective gets spent badly. An RTO breach halts operations. Staff sit idle, and the meter runs at whatever your hourly cost happens to be.

An RPO breach is quieter and more expensive to unpick. Lost transactions have to be identified before they can be recovered, and identifying them means reconciling against bank statements or customer complaints. Veeam's research across 1,200 organisations that suffered a successful attack found that only 57% of compromised data was ultimately recovered. The remainder becomes rework or a gap in your records you cannot explain.

Then there is the compliance dimension, which attaches almost entirely to the data side. UK GDPR Article 33 requires a notifiable personal data breach to be reported to the Information Commissioner's Office within 72 hours of becoming aware of it, and an availability breach, which is accidental loss of access to or destruction of personal data, counts. A long outage with clean data is a revenue problem for business continuity. A short outage with corrupted data can be a regulatory one.

Set recovery time objective vs recovery point objective independently, because the same system rarely needs both to be tight. A reporting warehouse can be offline for a day without anyone noticing, but losing a week of ingested data would be painful. A customer-facing booking engine is the reverse: every minute offline costs bookings, but a five-minute data gap is recoverable from confirmation emails.

Ready to bring your ideas to life?

Book a free 30-minute discovery call with our team — we'll understand your objectives and advise how bespoke software, MVP development, or an extended team can help your business.

Price tighter targets

Every hour you shave off an RTO buys you capacity you pay for whether or not you use it. A four-hour target for disaster recovery planning means documented restore procedures and support available during the window when the failure happens. A one-hour target means warm standby infrastructure and someone contractually obliged to respond at 3am. A target measured in minutes means running duplicate production capacity across zones or regions permanently.

The cloud pricing structure makes this explicit. Amazon's compute agreement offers a 99.99% Region-Level commitment only for instances deployed concurrently across two or more Availability Zones. Run a single instance, and you drop to a 99.5% Instance-Level commitment, which permits roughly 3.6 hours of unavailability per month before any credit is due. The stricter number is an architecture you fund.

Tightening RPO follows a different cost curve, driven by how often you capture data and where you put it:

  • Daily backups give you an RPO of up to 24 hours and cost little beyond storage.

  • Hourly or four-hourly snapshots narrow the gap considerably and add moderate storage and orchestration overhead.

  • Continuous replication to a secondary site pushes RPO towards seconds, and brings the operational cost of running and validating that replica indefinitely.

Near-zero on either measure of recovery time objective vs recovery point objective carries cost that rises far faster than the benefit. Replication also faithfully replicates corruption and ransomware encryption, which is why the National Cyber Security Centre's guidance insists on the 3-2-1 rule of three copies on two devices with one offsite, and on keeping at least one copy offline. An organisation with continuous replication and no immutable copy has bought a fast route to restoring bad data.

The British Library shows where the money actually goes when recovery has not been priced in advance. The October 2023 Rhysida ransomware attack forced a rebuild, and recovery has cost at least £7 million, around 40% of its unallocated reserves, with services still returning more than two years later.

Set defensible targets

Bold flat infographic on bright orange background featuring a shield with a checkmark and icons for business impact, risk acceptance, and supplier conversation.

A defensible target survives contact with a finance director. That means it comes from business impact, and it carries a named person who accepted the risk. The sequence below produces targets you can take to a supplier conversation.

Business continuity impact

Start disaster recovery planning by mapping how the cost of an outage escalates. The first thirty minutes cost nothing because orders queue. Hour four triggers contractual penalties. Day two triggers press coverage. Work out the shape of that curve for each critical process from your own revenue per hour and the fully loaded cost of the staff who would be sitting idle.

Then identify the manual workarounds you genuinely have. NCSC chief executive Richard Horne put the requirement plainly when writing to UK business leaders: "Organisations need to have a plan for how they would continue to operate without their IT, and rebuild that IT at pace, were an attack to get through." Paper fallbacks buy you hours, and those hours belong in your RTO. Business continuity planning that ignores what staff can do offline sets targets tighter than they need to be.

The point at which disruption becomes unacceptable is a judgement about harm. For a clinical or safety-related system it arrives quickly. For an internal expenses tool it arrives in a week. Write down where that line sits and why, because that reasoning is what your board is approving.

Tolerable data loss

Convert data loss into a time value by working from transaction volume. If your platform processes 400 orders an hour at peak and each lost order requires a customer call to reconstruct, the labour cost of a two-hour gap is straightforward to calculate. Reproducibility matters as much as volume: data that exists in a supplier's system or in a customer's inbox is recoverable, while data that only ever existed in your database is not.

Reconciliation effort is the cost people consistently underestimate. Financial records that fall out of step with payment provider records take days of skilled finance time to align, and the work cannot be delegated. Add any legal obligation to retain complete records, and any scenario where a data gap causes direct harm to a customer, such as a missing safeguarding note or a lost medication record.

Take the largest acceptable figure across those tests and express it as a time for recovery time objective vs recovery point objective. Two hours. Fifteen minutes. Twenty-four hours. That number is your RPO, and it should be recorded next to the reasoning that produced it.

Map critical dependencies

An application-level target for recovery time objective vs recovery point objective means nothing if the components underneath it recover more slowly. Your map needs to cover databases and the integrations that carry data in and out. It also needs to cover people, because a four-hour RTO that depends on one engineer who happens to be on annual leave is not a four-hour RTO.

October 2025 made the argument better than any diagram. A race condition in DynamoDB's automated DNS management left an empty record for a regional endpoint, and the resulting failure ran for more than 15 hours for some customers. It took down Slack and Atlassian along the way. As Cisco ThousandEyes noted, the incident showed "how a failure in a single, centralized service can ripple outwards through dependency chains that aren't always obvious from architecture diagrams."

Your slowest dependency sets your real recovery time. If a third-party payment integration takes six hours to reauthenticate after a failover, your two-hour RTO is fiction regardless of how quickly your own servers come back.

Ready to bring your ideas to life?

Book a free 30-minute discovery call with our team — we'll understand your objectives and advise how bespoke software, MVP development, or an extended team can help your business.

Approve target trade-offs

Assign two owners to every target for recovery time objective vs recovery point objective: someone in the business who owns the consequence of the outage, and someone technical who owns the delivery cost of meeting it. The business owner proposes the target based on impact. The technical owner prices it. Where the affordable capability falls short of the preferred target, the gap gets recorded.

That record is the point of the exercise. "We wanted one hour, and we can fund four" is a governance artefact once the Operations Director accepted the residual risk on 14 March. It also becomes the brief for the next investment cycle. The CEO of Evolved Ideas has spent over thirty years advising companies on their technical decisions, and her position on outsourced control applies directly here: "Our CEO Colette Wyatt recently raised concerns about the long-term risks of handing over control of your software to third parties," her team wrote after the Builder.ai collapse left customers locked out of their own platforms. Accepting a recovery target you don't understand is the same act of delegation.

RTO and RPO tiers

The ranges below are starting points for a conversation about recovery time objective vs recovery point objective. Adjust them against your own downtime cost curve and your sector's regulatory position.

Tier 1: high criticalityTier 2: medium criticalityTier 3: low criticality
Indicative RTO15 minutes to 2 hours4 to 12 hours24 to 72 hours
Indicative RPONear-zero to 15 minutes1 to 4 hours24 hours
Business impactRevenue stops, customer commitments breached, regulatory exposureOperations degrade, manual workarounds cover part of the loadInternal inconvenience, no external impact
Backup approachContinuous replication plus immutable copiesSnapshots every 1 to 4 hoursNightly backup
FailoverAutomated, multi-zone or multi-regionWarm standby, manual promotionRestore from backup
Support cover24/7 with contractual restoration timesExtended hours, defined responseBusiness hours
Testing cadenceQuarterly full failover testTwice yearly restore testAnnual restore test

Testing is where most tiering exercises quietly fail. Roughly 7% of organisations never test their disaster recovery planning at all, and of those that do, half test once a year or less. A Tier 1 designation with an annual tabletop exercise behind it is a label.

Complete the target worksheet

Run one of these per application before you approach a hosting provider or a support partner. The fields exist to force a conversation between the people who own the consequence and the people who own the build.

  1. Application name and business owner, with a named technical owner alongside.

  2. Critical business processes the application supports, and the hours during which each one matters.

  3. Estimated downtime cost per hour, split between lost revenue and lost productivity.

  4. Maximum tolerable outage before impact becomes unacceptable, with the reasoning recorded.

  5. Maximum tolerable data loss expressed in time, with the reconciliation effort behind that figure.

  6. Dependencies including database and named people.

  7. Current backup frequency and where copies are held, including whether one is immutable or offline.

  8. Failover method and whether it has been executed under real conditions.

  9. Third-party commitments that apply, and what they actually promise.

  10. Incident ownership and escalation path outside business hours.

  11. Date and evidence of the last successful recovery test.

  12. Approved RTO and approved RPO.

Field eleven does most of the work. A backup report showing successful completion tells you a file was written. Only a timed restore tells you whether you can meet the target you approved, which is why disaster recovery planning without test evidence is documentation.

Validate service commitments

With approved recovery time objective vs recovery point objective in hand, comparing providers becomes an evidence exercise. Three commitments get conflated in most proposals, and separating them changes what you are buying.

Response time is how quickly someone acknowledges your ticket. Availability is a statistical promise about uptime across a billing period, backed by a service credit. Restoration time is a commitment to have your service working again, and it is the only one that maps to your RTO. Ask specifically which one is being offered, because a 15-minute response commitment on a system with a 2-hour RTO leaves 105 minutes entirely uncovered by contract.

Check what the vendor service level agreement actually pays out. Under Amazon's Region-Level terms, monthly uptime below 99.99% but at or above 99.0% earns a 10% service credit against that service's monthly bill. If an hour of downtime costs you six figures, that credit is a rounding error. Treat availability percentages as a description of expected behaviour.

Then ask for recovery test evidence. European financial entities have had this codified since Article 11 of DORA required them to test continuity and disaster recovery planning at least yearly, which includes scenarios for switchovers between primary infrastructure and redundant capacity. That standard is a reasonable benchmark whether or not you're regulated. Ask when the provider last failed over and how long it took. A supplier who has never rehearsed the failover cannot commit to your restoration window, and business continuity depends on rehearsal.

Escalation paths for business continuity deserve the same scrutiny. Find out who has authority to declare a disaster and how you reach them at 2 am on a bank holiday. The FCA found that firms consistently underestimate this: its rules make clear it remains the firm's responsibility to stay within its impact tolerances even where a third party delivers the service.

Review resilience requirements

Quantified recovery time objective vs recovery point objective change what you buy. Once you know that a four-hour outage costs a specific amount and a two-hour data gap creates a specific reconciliation burden, architecture and supplier decisions become comparisons of cost against exposure. Test the targets and revisit them when the business changes.

Evolved Ideas builds and supports bespoke software for organisations across the UK and Europe, which means we spend a good deal of time working through recovery assumptions with clients before they harden into architecture. If you are weighing up recovery time objective vs recovery point objective for a business-critical system, speak to our team.

Ready to bring your ideas to life?

Book a free 30-minute discovery call with our team — we'll understand your objectives and advise how bespoke software, MVP development, or an extended team can help your business.

Yes, if those parts support distinct business processes and can recover independently. A booking site’s payment function can need a shorter target than its reporting area. Document the technical dependencies first, because a shared database or identity service can prevent separate recovery.

Calculate it from the longest period between usable recovery copies, including overnight and weekends. If a backup runs at 22:00 and fails at 08:00, the exposure is ten hours unless another copy captures the missing records. Check whether transactions can be reconstructed from source systems.

Record the shortfall, its business consequence, and the person who accepts it. Recovery time objective vs recovery point objective are approved risk limits, so an unfunded target should be revised or funded rather than left as an assumption. Set a date to review the decision after the next budget cycle.

Yes, include software-as-a-service tools if their loss stops a critical process or holds essential records. Your team may be unable to restore the provider’s platform, but it can plan for exports, alternate procedures, access recovery, and supplier escalation. Confirm what data the supplier retains and how it can be retrieved.

Ask Evolved Ideas to review targets before you select hosting, support cover, or failover architecture. A review can compare your approved outage and data-loss limits with dependency recovery times and test evidence. It also gives business and technical owners a shared record of assumptions.

Schedule a Discovery Call

Book a free 30-minute call that works for you

You Might Also Like

Discover more insights and articles

Over-the-shoulder view of hands typing on a laptop in a modern tech studio, with multiple screens displaying code and a coffee cup nearby.

Flutter vs React Native for B2B Mobile Apps: A Buyer’s Decision Guide

Picking a mobile framework by gut feel is how commercial trade-offs get made by accident. This guide translates the technical differences between the two leading options into terms that matter before you brief suppliers: cost, timeline risk, and talent availability. It includes a weighted scoring method you can run against your own requirements to reach a defensible decision rather than a default one.

A UK buyer evaluates app development teams in a modern workspace, featuring natural light, a premium desk, and minimalist decor.

App Developers Near Me: How UK Buyers Should Evaluate Local Fit

Picking a software partner because they're a short drive away feels safe, but proximity solves fewer problems than you might think. This guide separates what proximity genuinely improves from what it doesn't guarantee, and matches collaboration models (co-located, nearshore, extended team) to what your project actually needs.

A modern tech startup workspace with exposed brick, featuring diverse professionals collaborating around sleek devices and a central digital tablet.

AI Governance Framework for SMEs: Controls That Scale With Adoption

Most AI governance frameworks are built for enterprises with dedicated compliance teams. But what about lean teams that prioritise a fast pace? Here's an alternative: risk tiers to triage exposure, and a phased rollout sized for smaller organisations.

A modern tech startup office with a premium desk split into two zones, featuring a biometric login laptop and an access management dashboard.

Authentication vs Authorisation: What Business Software Teams Need to Get Right

Authentication and authorisation solve different problems: verifying who a user is and deciding what they can access. Understanding the difference helps you assess the cost and risk of getting either one wrong.