Backups give you copies to recover from. Disaster recovery gives your team a tested way to restore the services people need. To protect revenue and reduce disruption, define how quickly each critical service must return, how much data loss is acceptable and who will make recovery happen.

Person using a laptop displaying a cloud backup illustration

Backup protects copies. Recovery restores work.

Backup creates recoverable copies of data. Disaster recovery (DR) is the plan, infrastructure and coordinated work needed to restore IT services after a serious disruption. A backup is part of that plan, but a successful backup job does not tell you when staff can process orders again.

What backup and disaster recovery each need to prove
QuestionBackupDisaster recovery
What is protected?Recoverable data and system copies within the agreed scope.The services and dependencies needed to resume business activity.
What must be ready?Usable recovery points, retention, secure access and restore tools.Recovery capacity, an ordered plan, people, access and business checks.
What proves it works?A verified restore of the required data.A timed recovery in which users complete the required workflow.

Business continuity is broader again. It covers how people communicate, work around an outage and serve customers while technology is being restored. An SMB needs those practical arrangements alongside its technical recovery plan.

Set RTO and RPO for each important service

Start with two business decisions. They should guide your recovery design and be checked through testing.

RTO: how long can the service be unavailable?

The recovery time objective is the target recovery window. Define when the clock starts and what counts as an acceptable restored service. A server starting successfully is not enough if users still cannot sign in or complete a transaction. See NIST’s RTO definition.

RPO: how far back can the recovered data go?

The recovery point objective defines the point in time to which data must be recovered. It is commonly expressed as a maximum tolerable period of data loss. Backup frequency, replication lag and the availability of a usable recovery point affect whether you meet it. See NIST’s RPO definition.

Illustrative example: an order system fails at 10 a.m. A two-hour RTO targets usable service by noon. A 15-minute RPO requires a usable recovery point no older than 9:45 a.m. Those are separate objectives, and neither is a guarantee until the design has been tested under relevant conditions.

Agree targets with the service owner. A production application may need a one-hour recovery target while finance can tolerate a same-day restore. Those are example priorities, not universal requirements. Do not buy the same recovery tier for every workload by default.

Estimate downtime costs using your own business

Outage costs depend on what stops, when it stops and how long the disruption lasts. Review missed transactions, lost productive time, recovery labour, overtime, customer handling and contractual consequences. Distinguish sales that are delayed from sales that are permanently lost, and avoid counting the same loss twice.

Large headline figures need context. ITIC’s 2024 survey reported hourly downtime costs above $300,000 for more than 90% of midsize and large enterprises, with some reporting much higher costs. Those findings are not a default estimate for a small business.

For a Canadian perspective, Statistics Canada reported that total recovery spending following cyber security incidents rose from about CAD $600 million in 2021 to CAD $1.2 billion in 2023. The survey covered enterprises with at least 10 employees across most sectors. These historical totals show the scale of recovery spending, not the likely bill for an individual SMB.

Use your estimate to compare recovery options. A faster service target may justify extra capacity for the system that takes orders, while a less urgent archive may use a lower-cost restore approach.

Person holding a tablet beneath illustrated cloud, file, email and security icons

Protect the recovery copies and environment

Use layers that address different failures. Offsite copies help when a location is unavailable. Offline or appropriately immutable copies can help protect recovery data from deletion or alteration. Encryption protects confidentiality, but access to the required keys must also survive an incident.

The Canadian Centre for Cyber Security recommends preparing backups and a recovery plan as part of ransomware readiness. Design those protections around the incidents you need to withstand.

  • Separate access: protect backup administration so a compromised production account does not automatically control recovery copies.
  • Verify retention: confirm the available restore points and how long protected copies remain available.
  • Check recovery capacity: identify the infrastructure, storage, licensing and network access needed to run restored workloads.
  • Preserve essential material: keep recovery instructions, contacts, configuration records and secure access to keys available during an outage.
  • Include cloud workloads: confirm what your SaaS and cloud arrangements actually cover, including restore scope and limits.

Replication is not a complete backup strategy. It can help resume service quickly, but it may also copy unwanted changes. Retain recovery points and test which approach works for deletion, corruption, infrastructure failure and a security incident.

A separate cloud region can reduce some availability risks. It does not, by itself, prevent an attacker with shared administrative access from affecting the recovery environment. Review location and identity separation together.

Write a plan someone else can follow

1. Name the service owner and recovery lead

For each critical service, record the business owner, technical lead, escalation contacts and authority to declare a disaster. Include third-party support and the people who approve a return to service.

2. Map dependencies and recovery order

Identify identity services, networking, DNS, databases, storage, application servers and integrations. Identity and communication are often early priorities, but the order must follow the actual dependency map and incident conditions.

3. Document the recovery steps

Record the selected restore or failover method, prerequisites, verification checks and exceptions. Automate repeatable steps where appropriate, while keeping approval points for actions with business impact. Include credential rotation and device reconnection when the incident requires them.

4. Separate outage recovery from cyber recovery

A hardware failure and a suspected compromise need different decisions. In a cyber incident, coordinate containment and investigation, establish a suitable recovery point and check that the destination environment is ready before reconnecting services.

5. Plan the return to normal operations

Temporary recovery infrastructure still needs monitoring and support. Document how changes made during the outage will be reconciled and how the service returns to its normal environment. Tell staff which systems they should use at each stage.

Test the workflow, not just the restore button

A tabletop exercise checks decisions, contacts and handoffs. A technical restore test checks whether data and systems can be recovered. An end-to-end exercise checks whether staff can actually use the recovered service. Each answers a different question.

For critical services, consider quarterly recovery exercises as a starting schedule, then adjust the scope and frequency to risk, system changes and business constraints. Re-test after major changes. Plan exercises in a controlled environment with clear safeguards for production.

  • Record the scenario, scope, recovery point and start time.
  • Measure when the agreed business workflow becomes usable.
  • Check data freshness and record any gap against the RPO.
  • Validate identity, permissions, integrations and representative transactions.
  • Record workarounds, failed steps and dependencies that were missed.
  • Assign fixes to owners and set a date to verify them.

Veeam’s 2024 research, based on a commissioned survey of 1,200 IT leaders and implementors, highlighted concern about recovering critical data. It reinforces a useful question for your own review: what has your team demonstrated through a recovery test?

Person using a laptop with illustrated file transfers and a progress bar

Make provider responsibilities explicit

A managed service provider can help with monitoring, backup protection, recovery infrastructure, automation and testing. The scope varies. Confirm who responds after hours, which workloads are covered, what triggers failover and what is charged separately.

Ask for recent test results against your agreed targets, the open issues and the plan to resolve them. Recovery documentation can support audit or insurance discussions, but it does not guarantee coverage, favourable terms or a particular outcome.

ThinkSwift designs recovery plans for Canadian SMBs around the services that keep the business operating. That includes setting measurable recovery targets, planning protected offsite copies, assessing failover options and defining test reporting. Discuss the required scope and testing schedule through our technology services.

The useful starting point is one critical service: identify its owner, agree on acceptable downtime and data loss, then demonstrate its recovery.

Questions answered

01Do backups guarantee that we can recover?

No. Recovery depends on usable copies, the right restore point, working tools, required access and a suitable destination. A tested restore provides evidence that a specific recovery path works under the tested conditions.

02Do we need replication for every system?

No. Choose the recovery method against each service’s objectives. Some systems can meet their targets through backup restores; others may justify standby infrastructure or replication. Account for dependencies and cost.

03Are immutable backups enough for ransomware?

They can help preserve recovery copies against alteration or deletion during their retention period. They do not replace secure administration, access to recovery keys, incident response or testing, and they do not undo data theft.

04Should email and identity always come back first?

They are often essential to access and communication, but there is no universal order. Restore according to the dependency map, incident conditions and business priorities. Maintain an alternative communication method for the recovery team.

05What should a recovery test report show?

The scenario, systems tested, recovered data point, elapsed time to usable service, business checks, missed targets and corrective actions. Note any conditions that make the exercise different from a real incident.

Find out what recovery would look like for your business.

Talk to us about your critical services, current backups and the recovery targets your team needs to meet.

Talk it through
ThinkSwift Technology Team