Skip to content

Design a Backup and Recovery System

Published 2 October 2026

Likely problem statement

Design a backup and recovery system (think a simplified Backblaze, Druva, or the backup side of Rubrik). An organization has many machines, such as employee laptops, servers or VMs. Their data must be backed up to cloud storage on a schedule, and users or admins must be able to restore files, folders or whole machines to any earlier point in time.

Functional requirements

  • Register a machine or data source and set a backup policy: what to back up, how often (for example hourly or daily), and how long to keep it.
  • Take a full backup first, then incremental backups that upload only what changed.
  • Restore a single file, a folder, or an entire machine as of a chosen point in time.
  • Browse and search backup history ("show me this folder as it was last Tuesday").
  • Apply retention policies that delete old snapshots automatically, for example keeping 7 daily, 4 weekly and 12 monthly.
  • Show backup job status and alerts for failed or missed backups.

Non-functional requirements and scale

  • Around 10K–1M machines, each with 100 GB–1 TB of data, so petabyte-scale storage.
  • Durability above everything (11 nines). Losing backup data is the worst possible failure.
  • Backups must not slow down the source machine, so you need throttling and low CPU and network use.
  • Each backup must be consistent: a snapshot shouldn't hold a half-written file.
  • Targets for RPO (how much data you can afford to lose) and RTO (how fast you must restore).
  • Storage efficiency through deduplication and compression, since most data doesn't change between backups.
  • Encryption in transit and at rest, possibly client-side encryption with per-tenant keys.
  • Multi-tenant isolation.

Practical issues the interviewer probably raised

Since you said they kept bringing up real-world problems, these are the usual ones:

  • A 50 GB file changes by a few bytes. Do you re-upload all of it? This leads to splitting files into chunks (fixed-size or content-defined/rolling-hash chunking) and deduplicating chunks by hash.
  • The network drops halfway through a backup. Uploads need to be resumable and chunk uploads idempotent, and a snapshot is only "committed" once its manifest is written.
  • Files change while the backup is running. Use filesystem or VSS snapshots, or copy-on-write, to get a consistent view.
  • Deleting old snapshots when chunks are shared between them. This needs reference counting or a mark-and-sweep garbage collector, and you have to handle the race where a new backup starts using a chunk that GC is about to delete.
  • Ransomware encrypts the source, and the next backup uploads the garbage. Answers include immutable or WORM storage, anomaly detection on change rates, and retention that the client can't override.
  • Restoring 1 TB quickly. Options are parallel chunk fetches, restore prioritization, or shipping a physical disk.
  • Storage cost. Move old snapshots to cold tiers such as Glacier, and weigh that against slower restores.
  • Metadata scale. Storing snapshot manifests and file-to-chunk maps for billions of chunks, and deciding which database to use.
  • A storage region goes down. Replicate across regions, possibly with erasure coding.
  • Thousands of clients start backing up at midnight together. Add scheduling jitter and admission control.

Core design you'd have been steered toward

Components:

  • A client agent that scans files, chunks them, hashes the chunks and checks which ones are new.
  • A metadata service holding snapshots, manifests and the chunk index.
  • An object store for the chunks themselves.
  • A scheduler or policy service.
  • A GC and retention worker.
  • A restore service.

Backup data flow:

  1. The client takes a snapshot of the source.
  2. It splits files into chunks and hashes them.
  3. It asks the server which hashes are missing.
  4. It uploads only the missing chunks.
  5. It writes the manifest, which commits the snapshot.