An IT disaster recovery plan (IT DRP) is a documented, tested set of runbooks, roles and recovery targets that restore critical IT services within defined timeframes after an outage, breach or hardware failure. Per NIST’s definition, it is a written plan for recovering one or more information systems, and it sits inside the broader business continuity framework. If you do not have one yet, or yours has never been tested, start here.
Three actions to take right now:
- Quick inventory. List every server, workstation, cloud service and SaaS application your business depends on. Even a rough spreadsheet beats nothing.
- Identify your top one or two mission-critical systems. Ask: “If this goes down for 24 hours, can we still trade?” Those systems drive your first recovery targets.
- Confirm your most recent backup and your recovery contact list. When did the last backup complete? Who do you call at 2 AM when it fails?
Pro Tip: The first plan review must include business owners or senior managers, not only IT staff. They are the only people who can tell you which systems are genuinely mission-critical versus merely convenient.
Key takeaways
A tested IT disaster recovery plan, grounded in a business impact analysis with defined RTOs and RPOs, is the single most effective control an Australian SME can put in place to survive a ransomware attack or major outage.
| Point | Details |
|---|---|
| Start with the BIA | Recovery priorities and RTO/RPO targets must come from business impact analysis, not IT assumptions. |
| Test before you trust | Plans that have never been exercised fail at rates exceeding 70% in real incidents; schedule a tabletop before the plan is finalised. |
| Match backup type to RPO | A nightly full backup cannot meet a 4-hour RPO; use snapshots, replication or CDP for low-RPO systems. |
| SaaS workloads need dedicated backup | Microsoft 365 data requires a third-party backup tool with granular restore; relying on slow exports is not a recovery strategy. |
| Techbug’s approach | Techbug delivers vendor-agnostic DR planning, tested runbooks and emergency response for Australian SMEs, grounded in a real BIA. |
Table of Contents
- What does an IT disaster recovery plan actually protect?
- What must every IT disaster recovery plan contain?
- How do you build an IT disaster recovery plan step by step?
- What backup strategy actually meets your RPO?
- How do you set RTO and RPO for each system?
- How do you test a disaster recovery plan properly?
- What does building a DR plan actually cost?
- Which standards and templates should you use?
- How does your IT DR plan fit into the wider business continuity plan?
- How Techbug implements DR planning for Australian SMEs
- The mistakes that actually break DR plans (and the shortcuts that fix them)
- Techbug builds, tests and maintains DR plans for Australian businesses
- Sources
What does an IT disaster recovery plan actually protect?
An IT DRP protects your ability to keep trading when technology fails. It sits one level below the business continuity plan (BCP): the BCP defines which business functions must survive a disruption; the IT DRP defines how the technology that supports those functions gets restored. Ready is explicit that the two documents should be developed together, with IT recovery priorities derived from the business impact analysis (BIA).
For Australian businesses, the stakes are concrete. The Australian Cyber Security Centre (ACSC) consistently flags ransomware, business email compromise and supply-chain attacks as the top threats to SMEs. Data protection obligations under the Privacy Act 1988 and the Notifiable Data Breaches (NDB) scheme mean a prolonged outage that exposes personal information can trigger reporting requirements within 30 days of becoming aware of an eligible breach.
Common causes of IT disasters in Australian SMEs:
- Ransomware and malware (the ACSC’s Annual Cyber Threat Report names this the most disruptive threat to Australian organisations)
- Hardware failure (failed drives, power surges, ageing servers)
- Human error (accidental deletion, misconfigured cloud permissions, wrong firewall rule)
- Cloud region or SaaS provider outage (Microsoft 365 service degradation, AWS region failure)
- Supplier or third-party failure (ISP outage, managed service provider incident)
- Natural disaster (flood, fire, cyclone — all relevant in Queensland and northern Australia)
The numbers are sobering. Continuity Hub’s testing research reports that recovery plans which have never been exercised frequently fail in real incidents. For an Australian SME without a tested plan, a ransomware attack is not a question of whether the plan works — it is a question of whether there is a plan at all.
What must every IT disaster recovery plan contain?
A DR plan is only executable under pressure if it contains six core components: a current asset inventory, dependency maps, defined Recovery Time Objectives (RTOs) and Recovery Point Objectives (RPOs), documented runbooks, a roles and responsibilities matrix, and a communications and escalation procedure.

Asset inventory and dependency mapping
Start with hardware (servers, workstations, networking gear, UPS units), software (operating systems, line-of-business applications, licences), data (databases, file shares, email, SharePoint), credentials (admin accounts, API keys, MFA tokens) and third-party services (cloud platforms, ISPs, SaaS vendors). Dependency mapping goes one step further: it shows which systems must be recovered before others can function. Your accounting software cannot connect to the database if the database server is still offline.

A simple dependency diagram, even drawn in Lucidchart or draw.io, prevents the most common recovery mistake: restoring systems in the wrong order.
Runbook structure
Each runbook covers one system or service and should include:
- Scope: what the runbook covers and what it does not
- Prerequisites: systems that must be online first (dependency order)
- Step-by-step recovery procedure: numbered, specific, written for someone under stress
- Verification steps: how to confirm the system is actually working after recovery
- Escalation contacts: who to call if the procedure fails at step N
- Estimated time: realistic, based on a test, not a guess
Pro Tip: Write runbooks as if the person executing them has never touched that system before. Your lead sysadmin may be unavailable during the actual incident.
DR plan checklist
Use this as a starting point for your own downloadable checklist:
- [ ] Asset inventory completed and dated
- [ ] Dependency map created and reviewed
- [ ] RTO and RPO defined per system tier
- [ ] Runbooks written and version-controlled
- [ ] Roles and responsibilities assigned and accepted
- [ ] Communication and escalation tree documented
- [ ] Backup schedule confirmed and last restore tested
- [ ] Plan reviewed and signed off by executive sponsor
- [ ] Test schedule set (tabletop, component, full failover)
- [ ] Lessons-learned register created
How do you build an IT disaster recovery plan step by step?
Building a formal DR plan follows seven steps. Done in order, they prevent the most common failure: writing a plan that looks complete but cannot be executed because it was never grounded in actual business priorities.
-
Scope and stakeholder alignment. Define which systems, sites and services the plan covers. Get sign-off from an executive sponsor before investing further time. Without executive ownership, the plan will not be funded, tested or maintained.
-
Business impact analysis (BIA). For each business function, quantify the impact of downtime in financial, operational and reputational terms. The BIA output directly sets your recovery priorities. Ready.gov’s guidance is clear that IT recovery priorities must come from this analysis, not from IT’s own assumptions.
-
Risk assessment. Identify the threats most likely to affect your environment (ransomware, hardware failure, supplier outage) and score them by likelihood and impact. This shapes which recovery strategies you invest in.
-
Set RTO and RPO per system. Using BIA results, assign a Recovery Time Objective (maximum acceptable downtime) and Recovery Point Objective (maximum acceptable data loss) to each system tier. See the RTO/RPO section below for example mappings.
-
Select recovery strategies. Match each system tier to a recovery architecture: cold site, warm site, hot site, cloud failover, DRaaS, or a hybrid. Higher tiers demand more investment. Fortiv’s seven-step build guidance stresses that runbooks must be sequenced by dependency order, not alphabetically or by system name.
-
Write runbooks and assign roles. Document step-by-step recovery procedures for each system. Assign a named owner and a backup owner for every runbook. A roles and responsibilities matrix should cover: Incident Commander, IT Recovery Lead, Communications Owner, Executive Sponsor, and Business Process Owners.
-
Test and maintain. A plan that has never been tested is a hypothesis, not a capability. Schedule tests before the plan is finalised, not after. Review the plan at least annually and after any significant infrastructure change.
Typical development timeline:
| Organisation size | Discovery and BIA | Draft plan and runbooks | First test cycle | Total |
|---|---|---|---|---|
| Small SME (under 20 staff) | 1–2 weeks | 2–3 weeks | 1–2 weeks | 4–7 weeks |
| Medium SME (20–100 staff) | 2–4 weeks | 4–6 weeks | 2–4 weeks | 8–14 weeks |
| Larger SME (100–250 staff) | 4–6 weeks | 6–10 weeks | 4–6 weeks | 14–22 weeks |
What backup strategy actually meets your RPO?
The backup approach you choose must be matched to the RPO you set in the BIA. A system with a 4-hour RPO cannot rely on a nightly full backup. A system with a 24-hour RPO probably does not need continuous data protection.
Backup types and when to use them:
- Full backup: complete copy of all data; slow to run, fast to restore; suits weekly baseline for most systems
- Incremental backup: only changes since the last backup; fast to run, slower to restore (requires full + all incrementals); suits daily or hourly cycles
- Differential backup: changes since the last full backup; moderate run time, faster restore than incremental; a practical middle ground
- Snapshot: point-in-time copy at the storage or VM layer; near-instant, low RPO; suits virtualised workloads
- Replication: continuous or near-continuous copy to a secondary site or cloud region; lowest RPO; highest cost
- Continuous data protection (CDP): logs every write; RPO measured in seconds; used for databases and transaction systems
- SaaS backup (Microsoft 365, SharePoint Online): third-party tools that back up Exchange Online, SharePoint and Teams data independently of Microsoft’s own retention policies
On that last point: Microsoft’s own backup best practices whitepaper recommends planning for fast-scaled recovery rather than relying on slow export-based restores for Microsoft 365 workloads. When a ransomware event encrypts your SharePoint Online or Exchange Online data, a backup that takes weeks to restore from export is not a recovery strategy. Dedicated Microsoft 365 backup tools with granular item-level restore are the practical answer for backing up Office 365 emails and SharePoint Online data at speed.
Pro Tip: Test your backups by actually restoring from them, not by checking that the backup job completed. A backup that cannot be restored is not a backup.
Backup validation checklist
- Scheduled restore drills (at least quarterly for Tier 1 systems)
- Checksum or hash verification to confirm data integrity
- Encryption confirmed at rest and in transit, with key management documented
- Retention periods set and tested (including legal holds where required under Australian law)
- Offsite or cloud copy confirmed geographically separate from primary data
- Backup access credentials stored separately from production credentials
For CMS-driven businesses, platform-specific backup tools matter too. If you run a Joomla-based site, for example, dedicated Joomla backup plugins handle CMS-layer backups that generic file-level tools might miss.
How do you set RTO and RPO for each system?
The Recovery Time Objective (RTO) is the maximum time a system can be offline before the business impact becomes unacceptable. The Recovery Point Objective (RPO) is the maximum amount of data loss, measured in time, that the business can tolerate. Both figures come from the BIA, not from IT preference.
A practical way to derive them: for each system, ask the business owner “How long before this outage costs us real money or stops us trading?” That answer is your RTO ceiling. Then ask “How much data could we afford to re-enter manually if we lost it?” That answer sets your RPO ceiling.
Example RTO/RPO mapping by priority tier:
| Priority tier | Example systems | Target RTO | Target RPO | Likely strategy |
|---|---|---|---|---|
| Mission-critical | ERP, payment processing, core database | Under 1 hour | Under 1 hour | Hot standby, replication, DRaaS |
| Important | Email, file shares, CRM | 4–8 hours | 1–4 hours | Warm site, snapshot + cloud restore |
| Deferrable | Reporting tools, dev/test environments | 24 hours | 24 hours | Cold site, daily full backup |
Tightening RTO and RPO costs money. Moving from a 24-hour RTO to a 1-hour RTO for a core system typically requires replication infrastructure, a warm or hot standby, and tested failover automation. Document the cost of each tier and get explicit sign-off from the executive sponsor on the accepted thresholds. That sign-off protects IT when a recovery takes longer than a stakeholder expected.
How do you test a disaster recovery plan properly?
Testing converts a DR plan from a document into a capability. Continuity Hub’s research is direct: plans that have never been exercised fail at rates exceeding 70% in real incidents. The four test types, in order of complexity and disruption, are:
- Tabletop exercise: a facilitated discussion where the team walks through a scenario (e.g., “ransomware detected at 9 PM Friday”) and talks through decisions and actions. No systems are touched. Low risk, high value for testing communications and decision-making.
- Component test (walkthrough/simulation): a specific runbook or recovery procedure is executed in a non-production environment. Tests technical steps without risking live systems.
- Parallel test (sandbox failover): systems are brought up in an isolated environment alongside production. Validates that recovery procedures work without cutting over live traffic.
- Full interruption test (full failover): production is failed over to the DR environment. Highest fidelity, highest risk. Reserved for Tier 1 systems with mature plans.
N-able’s testing guidance recommends a progression through these types, with clear test objectives, safety controls and post-test remediation reports. AWS’s disaster recovery whitepaper adds a critical caution: assumption-based planning regularly hides failure points. Expired API keys, misconfigured firewall rules and cloud quota limits are invisible on paper but surface immediately in a real test.
Recommended cadence:
- Quarterly: tabletop exercise for Tier 1 and Tier 2 systems
- Semi-annual: component test for Tier 1 runbooks
- Annual: parallel or full failover test for Tier 1 systems
- After every major infrastructure change: targeted component test for affected systems
Pro Tip: Involve non-technical staff in tabletop exercises. TechTarget’s DR testing guidance notes that human error and communication breakdowns are among the most frequent causes of real-world plan failure. A test that only involves IT misses half the failure modes.
What to measure after each test
Capture actual RTO and RPO achieved (not the target), time from incident detection to plan activation, number of runbook deviations, and any steps that required improvisation. Feed every deviation back into the plan as a documented action item with an owner and a due date. A lessons-learned register that never gets actioned is just a list of known failures.

What does building a DR plan actually cost?
The biggest cost drivers are not always the obvious ones. Recovery architecture (replication, DRaaS licences, cloud storage) gets most of the attention, but staff time for BIA interviews, runbook authoring and test exercises often exceeds the infrastructure spend for SMEs.
Main cost drivers:
- Recovery architecture: replication tools, DRaaS subscriptions, cloud storage for offsite backups, hot/warm standby infrastructure
- Staff time: BIA interviews, runbook writing, test facilitation, post-test remediation
- Third-party DR services: managed DR providers, consulting fees for plan design or testing
- Data egress and replication costs: cloud providers charge for data transfer; high-volume replication adds up
- Testing disruption: full failover tests may require scheduled downtime and staff overtime
- Ongoing maintenance: plan reviews, annual tests, licence renewals
Indicative development timeline and budget tiers:
- Small SME (under 20 staff, simple environment): 4–7 weeks total; primary cost is staff time (20–40 hours) plus any new backup tooling. Expect $2,000–$8,000 AUD for an externally assisted plan and first test.
- Medium SME (20–100 staff, mixed on-prem and cloud): 8–14 weeks; staff time 60–120 hours; infrastructure changes and DRaaS licences add cost. Budget $10,000–$30,000 AUD for a full engagement including testing.
- Larger SME (100–250 staff, complex dependencies): 14–22 weeks; significant consulting and infrastructure investment. Costs vary widely based on recovery architecture choices.
These figures are indicative. The actual cost depends heavily on how mature your current backup and monitoring environment is before you start.
Which standards and templates should you use?
Three frameworks are most relevant for Australian organisations building a formal DR plan.
ACSC guidance is the Australian-first starting point. The Australian Cyber Security Centre publishes practical guidance on business continuity, backup strategies and the Essential Eight controls. For any Australian SME, ACSC guidance takes precedence over generic international templates because it reflects the threat environment and regulatory context your business actually operates in.
ISO 22301 is the international standard for business continuity management systems. It provides a framework for the BCP that the IT DRP sits inside. Certification is optional but the standard’s structure (context, leadership, planning, support, operation, evaluation, improvement) is a solid template for any formal plan.
NIST SP 800-34 is the US government’s contingency planning guide for federal information systems. It is widely used as a reference internationally and provides detailed guidance on BIA methodology, recovery strategy selection and test planning. The NIST DRP glossary entry ties DR planning explicitly to contingency and continuity frameworks.
Additional resources:
- Ready: a practical, free template framework that covers inventory, BIA, RTO/RPO and recovery strategies
- ACSC’s Small Business Cyber Security Guide: free, Australian-specific, covers backup and recovery basics
- Microsoft’s M365 Backup best practices whitepaper: essential reading if your business relies on Microsoft 365
Pro Tip: Choose vendor-agnostic templates that keep BIA results in a separate document from runbooks. When your infrastructure changes, you update the runbook without rewriting the BIA, and vice versa. Mixing them in one document is the fastest way to end up with an outdated plan.
How does your IT DR plan fit into the wider business continuity plan?
The IT DRP is a component of the BCP, not a replacement for it. The BCP defines which business processes must survive a disruption and at what minimum level. The IT DRP defines how the technology supporting those processes gets restored. If the two documents are written independently, IT recovery targets and business recovery targets will almost certainly conflict.
The flow works like this: the BIA identifies which business functions are most critical and sets maximum tolerable downtime for each. The BCP translates those into recovery priorities. The IT DRP then maps each priority to the systems that support it and sets RTO/RPO accordingly. An IT team that sets its own RTOs without BIA input will often over-invest in recovering systems that are not business-critical and under-invest in ones that are.
Business roles that must be involved when the plan is activated:
- Executive sponsor: authorises the declaration of a disaster and the activation of the DR plan
- Business process owners: confirm which functions are impaired and validate recovery priorities in real time
- Communications owner: manages internal and external messaging (staff, customers, regulators, media)
- IT recovery lead: coordinates technical recovery against the runbooks
- Legal/compliance contact: advises on NDB scheme reporting obligations and any contractual notification requirements
For a deeper look at aligning IT recovery to business continuity, Techbug’s practical BCP guide for SMEs covers the BCP structure that the IT DRP feeds into.
How Techbug implements DR planning for Australian SMEs
Techbug’s approach to IT disaster recovery is vendor-agnostic, business-priority led, and built around tested runbooks rather than theoretical frameworks. With over 30 years of combined experience supporting Australian SMEs, the team starts every engagement by understanding the business before touching the technology.
Typical engagement phases:
- Discovery: audit of current infrastructure, backups, cloud services and existing documentation
- Business impact analysis: structured interviews with business owners and process owners to derive RTO/RPO targets
- Recovery strategy design: vendor-agnostic selection of backup tools, replication, DRaaS or cloud failover options matched to each system tier
- Runbook authoring: step-by-step recovery procedures written for each critical system, sequenced by dependency order
- Test exercises: facilitated tabletop exercises and component tests, with full failover testing for Tier 1 systems
- Ongoing maintenance: annual plan reviews, quarterly backup validation, and emergency response support when incidents occur
For Brisbane-based businesses that need hands-on data recovery support during an active incident, Techbug’s emergency response team is available to assist on-site.
Pro Tip: Ask any DR service provider to show you a test report from a previous engagement, not just a sample plan document. A tested plan with documented results is worth ten untested ones.
The mistakes that actually break DR plans (and the shortcuts that fix them)
Most DR plan failures come down to a handful of recurring mistakes, not exotic architecture problems. Experienced practitioners see the same patterns repeatedly.
Common mistakes:
- Assumption-based planning: writing “failover to cloud” in the plan without ever testing whether the cloud environment is actually configured, sized and accessible. AWS’s own guidance warns that unvalidated recovery paths regularly hide expired credentials, quota limits and firewall mismatches.
- Runbooks with no real detail: a runbook that says “restore the database from backup” is not a runbook. It is a reminder that a runbook needs to be written.
- Not testing high-risk dependencies: the systems most likely to cause a cascading failure (Active Directory, DNS, core databases) are often the ones least frequently tested because they feel too risky to touch.
- Unclear activation authority: if nobody knows who can officially declare a disaster and activate the plan, the first 30 minutes of a real incident are spent in a meeting instead of in recovery.
- IT-only ownership: a plan that only IT staff know about will fail the moment a communication or business-process decision is needed under pressure.
High-leverage shortcuts that deliver fast improvement:
- Tier your systems immediately. Even a rough three-tier classification (mission-critical, important, deferrable) lets you focus effort where it matters and stop treating every system as equally urgent.
- Pre-draft your communications. Write the staff notification, the customer notification and the regulator notification templates now, before an incident. Filling in a template under pressure takes minutes; writing one from scratch takes hours.
- Store credential fallbacks offline. Admin passwords, MFA recovery codes and API keys stored only in a cloud password manager are inaccessible if the cloud environment is the thing that failed. Keep an encrypted offline copy in a physically secure location.
- Run a tabletop exercise before the plan is finished. Walking through a scenario with the team will surface gaps faster than any document review.
Techbug builds, tests and maintains DR plans for Australian businesses
Techbug delivers vendor-agnostic DR planning, testing and emergency response for Australian SMEs, with a focus on Brisbane and Queensland businesses. The difference from a generic IT provider: every plan is grounded in a real BIA, every runbook is tested before it is signed off, and every backup is validated by restore, not just by job completion.

Services cover the full DR lifecycle: discovery and BIA facilitation, recovery strategy design, runbook authoring, automated backup validation, tabletop and failover testing, and ongoing plan maintenance under a managed IT services subscription. For businesses that need to address the cybersecurity side of DR, Techbug’s IT security services include ransomware-safe backups and ACSC Essential Eight implementation.
Ready to build a plan that actually works under pressure? Talk to Techbug about a DR planning engagement tailored to your business size, systems and compliance obligations.
Sources
- Ready
- disaster recovery plan (DRP) – Glossary | CSRC
- Microsoft 365 Backup best practices whitepaper
- Disaster Recovery Testing: Validation Frameworks, Automated Testing, and Exercise Design – Continuity Hub
