Home›Guides›DRP and BCP

DRP and BCP

What is an RTO?

The RTO (Recovery Time Objective) is the maximum time a service can remain unavailable. It is measured from the incident, or from the decision to fail over, until a user can once again carry out a normal business task. Not until a machine is powered on whose application has not yet been checked.

Updated October 20263 min read4 sources cited

Key points

  • The RTO adds up six delays: detection, decision, finding the access credentials, technical time, business verification, users getting back online.
  • NIST distinguishes it from the maximum tolerable downtime (MTD): the RTO should normally be shorter than the MTD.
  • An RTO is written per service: the switchboard and the archives do not have the same one.
  • Only a timed test tells you whether the written RTO is being met.
  • Support hours and the absence of on-call staff are part of the actual RTO.

Official definition

NIST defines the RTO as the maximum amount of time that an information system resource can remain unavailable before there is an unacceptable impact on the activities it supports. It distinguishes it from the maximum tolerable downtime (MTD), which is the total downtime management accepts for an activity. The RTO must ensure that the MTD is not exceeded: it is therefore normally shorter.

ANSSI, France’s national cybersecurity agency, uses the term maximum tolerable interruption time (DMIA). It requires a backup strategy to take this into account for each business asset, and a restoration order to be defined in advance, based on dependencies (DNS, directory, etc.) and the criticality of the applications.

What the RTO is the sum of

For a conventional restore:

  1. the time to notice the failure;
  2. the time to decide and reach the person who knows what to do;
  3. the time to find keys, passwords and the procedure;
  4. the technical time for copying or starting up;
  5. the time for someone from the business to check;
  6. the time for workstations or remote customers to get the service back (DNS, VPN, IP).

A “two-hour” RTO advertised by a software vendor often counts only step 4, under lab conditions. The actual RTO adds up all six. At night and at weekends, step 2 alone can exceed two hours if nobody is on call.

For a BCP, steps 4 and 6 are prepared in advance. What remains is detection and the risk of a failover that nobody dares to approve.

RTO and RPO are not traded off against each other

You can have a short RPO (frequent copies) and a long RTO (slow restore of a large volume). You can have a short RTO (standby system already running) and a poor RPO if the standby system is two hours behind. Both figures must be written down.

RPORTO
Question askedHow much work can we lose?How long can we stay down?
MeasuredBackwards, from the incidentForwards, from the incident
Controlled byThe frequency of copiesThe preparation of the standby system
Verified byThe date of the last successful copyA timed test

One RTO per service

The switchboard and the archive document management system do not have the same RTO. Writing “RTO 4 hours” for the whole company forces you either to overpay for the archive system or to be untruthful about the switchboard. One line per service is enough.

How to know whether the RTO is met

Only with a stopwatch during a test. If the test took six hours and the written RTO is two hours, it is the written RTO that is wrong, until the architecture changes. You do not “aim” for an RTO that the last measurement has contradicted. ANSSI insists on this point: a restore procedure must be written and regularly carried out. The pace of testing is discussed in How often should you test your DRP?.

At WeDoBack

No single RTO figure is published on the site, and inventing one would be misleading: it depends on the volume, the link, the instance size and the availability of people on the customer side. What the architecture changes is the nature of the delay. With a simple restore, the data has to be brought back and possibly reinstalled. With the DRP, servers restart on standby instances from the chosen version: the technical delay is that of this restart, not that of buying a server. A boot test takes place every month, without touching production. With the BCP, cloud instances run permanently and are relayed by an agent on the customer’s network, without changing IP address: the residual RTO is mainly that of detection and decision. Replication or synchronisation of data between the BCP instance and the original server is not native: it requires a specific process, tailored to the need, which WeDoBack can set up on the basis of a quote. In all three cases, business verification remains part of the timing. Human support is available from 9 am to 1 pm and from 2 pm to 5:30 pm (Paris time, i.e. from 11 am to 3 pm and from 4 pm to 7:30 pm Mauritius time during European summer time, one hour later during European winter time).

Frequently asked questions

What is the difference between RTO and MTD?

The MTD (Maximum Tolerable Downtime) is the total downtime that management accepts for an activity, all impacts included. The RTO is the time needed to bring an IT resource back into service. NIST specifies that the RTO should normally be shorter than the MTD, to leave room for the other recovery steps.

A software vendor advertises an RTO of a few minutes. Is that realistic?

That figure generally covers only the technical start-up time, in a lab. It includes neither detection, nor the time to reach the authorised person, nor verification by a user. Your actual RTO is the one you measured during your last test, from the declaration of the incident to the first successful business task.

Is the RTO a legal requirement?

No text imposes a specific duration on an SME. However, the Data Protection Act 2017 (section 31) requires appropriate security measures against the loss or destruction of personal data, which implies being able to make it available again within a reasonable time after an incident. The RTO is the practical way of defining that time.

Planning a backup, DRP or BCP project?

More than 20 years of experience protecting business data.

Request a quote+33 9 72 50 78 28

Protect your data with WeDoBack

Encrypted offsite backup, immutable storage, DRP and BCP: tell us about your servers and we will recommend the right combination.