Tenant inventory
MSP backup operations break down when the team does not have a current picture of what each customer expects to recover. A backup console may show jobs and protected machines, but customer recovery depends on ownership, priority, retention, dependencies, and the last time a restore was actually proven.
The tenant inventory should be written for operations, not procurement. It should tell an engineer what systems matter, what policy applies, where recovery evidence lives, who can approve a restore, and which systems are intentionally excluded. Exclusions need explicit customer approval; otherwise they become surprises during incidents.
Tenant boundaries matter. Customer-specific runbooks, access paths, storage targets, and reports should be separated clearly enough that an incident or operational mistake in one customer environment does not leak into another.
- Maintain one protected-systems register per customer.
- Record workload owner, recovery priority, retention policy, storage target, and last verified restore.
- Mark systems that are excluded from backup with an explicit customer-approved reason.
- Keep customer runbooks separate enough that one tenant incident does not expose another tenant.
Policy baseline
A policy baseline keeps every customer from becoming a custom backup snowflake. The goal is not to force every workload into the same schedule. The goal is to make exceptions visible, reviewed, and tied to business need.
Useful baselines usually separate common workload classes: ordinary servers, databases, high-change systems, archived systems, and systems with special compliance or residency requirements. Each class should have a default retention window, storage target, verification expectation, and escalation path.
When a customer asks for an exception, the exception should be recorded in plain language. "Lower retention because the workload is rebuilt from source" is useful. Silent drift in backup policy is not.
- Define standard policies for servers, databases, high-change workloads, and archived systems.
- Document retention, immutability, off-site copy, and encryption expectations.
- Require named approval for exceptions instead of burying them in ticket comments.
- Review policy fit after migrations, new compliance requirements, or storage changes.
Daily exception review
Daily review should focus on recovery risk, not ticket volume. A failed job on a low-priority lab server and a missed backup on a customer identity server should not receive the same attention just because both are red.
The review should separate transient noise from patterns. A single retry may be harmless. Repeated failures, capacity drift, unexpected backup growth, or stale off-site synchronization can quietly remove recovery options while dashboards still look mostly healthy.
Every daily review needs a decision: ignore with reason, retry, escalate, change policy, or contact the customer. If the team cannot make that decision from the available data, the operating model needs better evidence.
- Review failed, missed, delayed, and unusually large backup jobs.
- Separate noisy transient failures from issues that reduce recovery confidence.
- Escalate repeated failures by workload priority, not by ticket age alone.
- Confirm that repository capacity and object-storage synchronization are not drifting silently.
Weekly restore verification
Weekly restore verification is where an MSP proves that backups are more than completed jobs. The sample does not need to include every system every week, but it should rotate through representative workload types and priority tiers.
A good verification record says which restore point was tested, where it was restored, what health checks passed, how long it took, and what follow-up is open. That record is useful for operators, customer reviews, audits, and post-incident evidence.
Failed verification should not disappear into a generic failure queue. It should become a recovery-risk item with customer visibility when appropriate, because the issue is not a backup inconvenience; it is a possible service outage later.
- Select representative systems across customers and workload types.
- Verify restore points, boot readiness, and application checks where possible.
- Record evidence that a customer or auditor can understand later.
- Turn failed verification into a dated remediation item with an owner.
Monthly customer review
A monthly backup review should not be a screenshot of green jobs. Customers need to understand whether their important systems can be recovered, what was tested, what changed, and which risks still need a business decision.
Use plain recovery language. "Last verified restore" is more meaningful than platform-specific job status. "Open risk: database restore not tested after migration" is more useful than a vague warning count.
The review is also where customer priorities get corrected. New systems, new contracts, new compliance expectations, and new revenue dependencies should change recovery order and backup policy. If the review does not ask about those changes, the backup plan slowly becomes stale.
- Report protected systems, last backup, last verified restore, open risks, and policy exceptions.
- Use plain recovery language instead of backup-platform internals.
- Separate technical success from business readiness.
- Ask whether recovery priority changed because of new systems, contracts, or operations.
Incident handoff
During an incident, the MSP should hand over recovery choices in a way the customer can act on. The customer does not need internal tooling detail; they need workload status, candidate restore points, confidence level, known limitations, and the decision needed next.
The handoff should distinguish backup availability from recovery readiness. A restore point may exist but still require dependency recovery, isolation, application validation, or security review before the service can return.
After the incident, preserve the recovery evidence. The next drill, review, or renewal conversation should start from what actually happened, not from generic service descriptions.
- Keep a named incident owner for backup recovery decisions.
- Share candidate restore points with confidence notes and known limitations.
- Track restore status by workload, dependency, and business service.
- Preserve the after-action record so the next drill starts from evidence.
What a useful customer report includes
- Protected systems and priority tier
- Most recent successful backup
- Most recent verified restore
- Retention and off-site copy status
- Open recovery risks and owners
- Changes requested from the customer
Turn this into a restore check
Use this operating rhythm to make customer backup reporting about recoverability, not only successful jobs.
Resource contents
Use this resource for
Planning, review, and evaluation. The content stays focused on recovery decisions and evidence, not proprietary implementation details.