Sendabrief logoSendabrief
Free template

Backup and restore verification

Almost every organisation backs up and almost none verify that a restore works. The failure mode is always the same: the backup job reports success for months, and the first real restore attempt reveals it was backing up the wrong thing, or that nobody knows the decryption key. A backup you have never restored from is a hypothesis, not a backup.

Download as Word (.docx) 11 stepsNo email, no signup. 11 steps, editable in Word or Google Docs.
Trigger

Daily for the automated job, and monthly for the restore test.

Roles involved
IT AdministratorSystem OwnerIT Lead
Review cadence

Quarterly, and immediately after adding any system, changing backup tooling, or any failed restore.

  1. 1

    IT Administrator confirms the backup job completed and checks the log rather than only the status indicator.

    IT Administrator

    Timing: Daily

    A job can report success while silently skipping files it could not read. The log shows what was actually captured.

    Completed cleanly: go to step 2

    Failed or partial: go to step 8

  2. 2

    IT Administrator confirms the backup size is within the expected range for the data set.

    IT Administrator

    A backup suddenly much smaller than usual is the classic sign of a source path that moved. The job still says success.

  3. 3

    IT Administrator confirms coverage against the list of systems and data that must be backed up.

    IT Administrator

    Coverage against a written list, reviewed when systems are added. New systems not added to the backup scope are the most common gap, and nobody notices until the restore.

    All in scope covered: go to step 4

    Gap found: go to step 9

  4. 4

    IT Administrator confirms at least one copy is held off-site or in a separate account, and is not writable from the primary environment.

    IT Administrator

    Separation is what makes a backup survive ransomware or a compromised admin account. A backup reachable with the same credentials as production is not a backup.

  5. 5

    IT Administrator performs a restore test to an isolated location, restoring real data rather than checking the job would run.

    IT Administrator

    Timing: Monthly

    Actually restore, to somewhere isolated. This is the only step that proves anything, and it is the one that gets deferred indefinitely.

    Restore succeeded: go to step 6

    Restore failed: go to step 10

  6. 6

    System Owner confirms the restored data is complete and usable, opening records and checking the most recent entries.

    System Owner

    The system owner checks usability, not IT. A technically successful restore of subtly corrupted data still fails the business.

    Usable: go to step 7

    Incomplete or corrupt: go to step 10

  7. 7

    IT Administrator records the test: date, what was restored, how long it took, and who verified it.

    IT Administrator

    Record the elapsed time. That figure is your real recovery time, and it is usually much longer than anyone assumes.

    Recorded: the procedure ends

  8. 8

    IT Administrator investigates the failure, reruns the job, and escalates to the IT Lead if it fails again.

    IT Administrator

    Two consecutive failures is an incident, not a task. Silent repeated failure is how organisations end up with no usable backup at all.

    Rerun succeeded: go to step 2

    Failed again: go to step 11

  9. 9

    IT Administrator adds the missing system to the backup scope and records why it was missing.

    IT Administrator

    Record the cause. It is nearly always a provisioning process that does not include adding new systems to backup, which is the real fix.

    Scope corrected: go to step 4

  10. 10

    IT Lead treats the failed restore as an incident, identifies the cause, and sets a date for a re-test.

    IT Lead

    A failed restore is the most serious finding this procedure can produce, because it means the protection you believed in does not exist.

    Fixed, re-test scheduled: go to step 5

  11. 11

    IT Lead escalates repeated backup failure, arranges an interim protection method, and informs whoever owns the risk.

    IT Lead

    Somebody outside IT needs to know the organisation is currently unprotected. That is a business decision, not a technical one.

    Interim protection in place: go to step 8

Change these before you use it

  • List the systems and data in scope explicitly, and add reviewing that list to your system provisioning procedure.
  • Set your recovery point and recovery time objectives, then check whether the elapsed time recorded in step 7 actually meets them.
  • Set the restore test frequency by criticality rather than uniformly monthly.
  • Confirm where decryption keys and credentials are held and that they are not stored only inside the system being backed up. This is a genuinely common single point of failure.
  • Add your retention periods, including any legal minimum, and confirm them with whoever owns your compliance.

This is a starting point, not compliance advice. It is written to be adapted, and a procedure that touches access, money, or customer data needs to match how your business actually operates and whatever rules apply to you. Use it as a first draft to edit, not a policy to adopt.

Make it yours in a couple of minutes

Rather than retyping this and editing it, describe your own version of the process out loud or click through it once, and get a first draft with your actual steps, roles, and systems in it. No account needed to see the result.

No account neededNo credit cardSee the whole SOP before you sign up