Recovering IT systems after an outage starts not with repairs but with two numbers: RTO, the time within which a system has to be back in service, and RPO, the amount of data the business is prepared to lose. They drive the backup scheme, server redundancy and the sequence of actions that is written down in advance rather than during the incident.
Key takeaways
The causes of incidents repeat year after year, and almost all of them are predictable. What separates companies is not whether outages happen, but how long the downtime lasts and how much is recovered.
Each cause has its own countermeasures, but they share one thing: you act on a plan written in advance, not on the memory of whoever noticed the problem first.
RTO and RPO turn a conversation about reliability into an engineering task. In the NIST definitions from SP 800-34 Rev. 1, the recovery time objective is the overall length of time a system's components can stay in the recovery phase before the organisation's work is unacceptably affected, and the recovery point objective is the point in time to which data must be recovered after an outage.
In practice these are two questions: how many hours the business survives without the system, and how large a period of data it can afford to lose. The answers are set by the manager rather than by the IT team: this is a decision about acceptable losses, and it defines the budget.
| System | What stops during downtime | What is lost without a fresh backup | What sets the numbers |
|---|---|---|---|
| Order intake and payments | Revenue: the customer cannot place and pay for an order | Payments and orders never written into the system | Orders per hour and the share of customers who will not come back |
| CRM | The sales team: calls, correspondence, deals in progress | Conversation history and agreements for the period | Deals in the pipeline and how quickly a client goes elsewhere |
| Accounting and inventory system | Shipments, goods receipt, closing documents | Ledger entries and stock movements for the period | The shipping schedule and reporting deadlines |
| Corporate mail | Client correspondence, confirmations, invoices | Incoming mail for the period | The share of agreements that live only in email |
| Website and client portal | Enquiries, self-service, advertising into a void | Content changes and enquiries that never reached the CRM | The share of enquiries that arrive through the site |
Once the numbers are set, it becomes clear what to pay for. A short RTO calls for spare capacity and a rehearsed switchover, a short RPO for frequent copies and replication. The reverse holds too: systems whose downtime can wait until the next working day do not need an expensive scheme.
The 3-2-1 rule is the baseline backup scheme: three copies of the data, on two different types of storage, with one copy kept off site. The CISA guidance for small and medium businesses states it as three copies of important files, two different types of storage media — a hard drive and the cloud, for instance — and one copy stored away from the business location.
Each number covers its own kind of disaster, and together they provide the margin.
A copy that cannot be reached from the working network deserves separate mention: disconnected from the network or protected from modification. The same CISA guidance advises testing restores both fully and partially and making sure data can be rolled back at least seven days.
Important A backup that has never been restored is a hypothesis, not protection. A test means deploying the data into a test environment and measuring how long it takes, not reading a log entry saying the job completed.
The most expensive mistake here is treating something as a backup when it is not one. All four technologies below are useful, but they protect against different things.
A backup differs from all of the above in that it is a separate copy of the data, fixed at the moment it was made, from which you can return to the state at a chosen point in time.
A recovery plan answers the questions before they turn into panic: what is fixed first, who makes decisions, where the backups are and how to reach the team if email is down. It is a working document of a few pages, not a volume for the shelf.
The plan is verified by drills: pick a system, restore it from a backup into a test environment and measure how long it actually took. The gap against the assigned RTO is the result the drill exists for.
What happens in the first hour decides the scale of the damage and whether any evidence survives for the investigation. The general sequence looks like this.
Communication with customers deserves its own line. Silence during downtime costs more than the downtime itself, so the wording of the notice is prepared in advance and kept next to the plan.
In a hardware failure the data is intact and the task comes down to moving it onto working equipment. In a ransomware attack the data is damaged deliberately, and the infrastructure is treated as compromised until proven otherwise — you cannot restore on top of it.
According to the Verizon Data Breach Investigations Report 2026, ransomware is involved in 48% of the breaches analysed, while companies increasingly choose not to pay and the amounts demanded are falling. Paying does not bring the data back by itself: it only shows attackers that the company is willing to pay.
How to build protection against such scenarios into systems at the design stage — encryption, network separation, logging — is covered in the article on security at the design stage.
An argument about the redundancy budget stays an argument about taste until downtime has a price. The calculation is simple: multiply the hours of downtime by the revenue lost and the wages of staff who cannot work, then add the one-off cost of recovery and outside help.
The resulting figure is compared with the cost of the measures: a spare server, a second site, more frequent backups, a support contract. If an hour of downtime costs more than a month of maintenance, there is nothing left to discuss. The data for such a calculation comes from an infrastructure review — what that work involves is covered in the article on the IT infrastructure audit.
Where the backups physically sit is a separate question. Where databases and copies are stored, who can reach them and what the law says about it is covered in the article on choosing between cloud and a server in Uzbekistan.
Syntra Systems starts working with someone else's infrastructure by taking it over: we describe the systems and their criticality, collect access, check backups and monitoring and build a risk map. A restore test is part of that handover, because this is where the gap between job reports and reality usually shows up.
After that the infrastructure lives by a regulation: monitoring with alerts, scheduled backups with regular restore tests, incident reviews down to the cause and a monthly report. Support of a company's business systems starts from 700 $ per month, website support from 250 $ per month; the scope is described on the technical support page. What regular maintenance includes and what is fixed in the contract is covered in the article on technical support for IT infrastructure.
When the plan calls for a second site or for moving systems closer to customers, that is separate work: migrating infrastructure without downtime starts from 900 $, and servers for heavy systems with redundancy and tested backups from 150 $ per month — terms are on the servers and cloud page.
Let’s discuss your project
Tell us what you need, and we will estimate the timeline and cost and suggest a solution.
As the only measure, no. A drive permanently attached to a computer gets encrypted along with it, and a copy in the same room will not survive a fire or a theft. The 3-2-1 rule calls for a second type of storage and one copy kept off site.
Restore it in a test environment and time how long it took — a log entry saying ‘job completed successfully’ doesn't count as a check. CISA recommends testing both full and partial restores and confirming data can be rolled back at least seven days.
Paying does not bring the data back by itself and shows attackers that the company is willing to pay. The working strategy is different: restore from a copy the malware could not reach and close the route it used to get into the network.
Yes, but a short one. A few pages are enough: the list of systems with priorities, where the backups are, who makes decisions, contacts outside corporate email and the restore procedure. The value of a plan is not its length but the fact that it has been tested at least once.
Isolate the affected machines from the network without powering them down unless there is no other way, record the picture and preserve the logs. Then restore from a copy made before the infection onto a freshly built environment and rotate every credential that could have reached the attacker.
The Law on Cybersecurity ZRU-764 obliges cybersecurity subjects to notify the authorised state body about incidents and to preserve digital evidence. Requirements for critical information infrastructure are stricter, so the reporting procedure is written into the plan in advance.
The head of the company, not IT — this is a decision about acceptable losses, and it drives the budget. The calculation multiplies downtime hours by lost revenue and the pay of staff who can't work, then compares that against the cost of a backup server or more frequent copies.