Live · Sun, Oct 4, 2026 · 07:01 UTC Block 843,917 Fees 14 sat/vB Fear & Greed 72 · Greed
Newsletter Pro Terminal Sign in
ITop Field News.
Subscribe →
Live · 07:01 UTC Block 843,917 F&G 72
Cloud & infrastructure Cloud & infrastructure desk

Cloud warm standby vs hot standby: which DR model fits your workload?

Warm standby and hot standby are often confused in DR planning, but they carry fundamentally different cost and recovery time implications. Here's how Australian IT teams should choose between them.

Close-up of server racks in a data center highlighting modern technology infrastructure.

Photo by panumas nikhomkhai on Pexels

Cloud disaster recovery planning in Australia often stalls on a single decision: how much availability can you actually afford, and what does "standby" really mean? The terms warm standby and hot standby both appear in architecture conversations and vendor documentation, but they describe recovery postures that are genuinely far apart in cost, complexity, and recovery time. Picking the wrong model is one of the quietest budget mistakes in cloud infrastructure.

What warm standby actually means

A warm standby environment runs a reduced but live copy of your production stack in a secondary region or availability zone. The databases are replicated and current. The application tier is running but at lower capacity, typically a fraction of production scale. When a failure hits, you promote the standby and scale up. That scale-up step is what separates warm from hot.

Recovery times for warm standby typically sit in the range of 10 to 30 minutes, depending on how aggressively you've pre-configured auto-scaling and how complex your promotion scripts are. That's fast enough for most enterprise workloads, but not fast enough for systems where every minute of downtime translates to direct revenue loss or safety risk.

The cost advantage is real. Running a warm standby at 20–25% of production capacity costs roughly 30–40% of what a full hot standby would. For an Australian organisation running mid-tier production workloads on AWS Sydney or Azure East Australia, that difference can be tens of thousands of dollars per month.

What hot standby actually means

A hot standby runs a full-capacity, fully synchronised mirror of production. Traffic routing switches in seconds, typically via DNS failover or a global load balancer. There's no scale-up step. The environment is already warm, fully provisioned, and actively processing health checks.

Recovery time objectives (RTOs) for hot standby can reach under two minutes. Recovery point objectives (RPOs) can approach zero with synchronous database replication, though that introduces its own latency trade-off on writes. AWS Route 53, Azure Traffic Manager, and Google Cloud Global Load Balancer all support the automated health-check-driven failover that hot standby depends on.

The cost is the constraint. A hot standby doubles your compute and data transfer footprint. It also doubles your licensed software costs on anything not covered by cloud-native tooling. For Australian organisations subject to data residency requirements, running a full hot standby in a second Australian region (say, AWS Melbourne alongside AWS Sydney) means paying for two complete regional deployments.

The four questions that decide your model

Before choosing, four variables matter most. First: what is your RTO tolerance? If your SLA permits 20 minutes of downtime, warm standby almost certainly meets it. If your SLA requires sub-five-minute recovery, you need hot standby. Second: what is your RPO tolerance? Warm standby with asynchronous replication can introduce a small data lag, typically under one minute. Hot standby with synchronous replication eliminates that lag but adds write latency.

Third: what are your actual cost constraints? This is where cloud bills in Australian businesses tend to go wrong in DR planning: teams spec a hot standby for every workload because it sounds safer, then carry that cost indefinitely on systems that haven't had an outage in three years. Fourth: what are your compliance obligations? The Australian Prudential Regulation Authority (APRA) CPS 230 standard, which governs operational risk for financial services entities, sets recovery expectations that can push regulated entities toward hot standby for core processing systems.

Workload tiering: the practical way to decide

Most Australian organisations don't have a single DR model. They have a tiered architecture where workloads sit in one of three categories: mission-critical (hot standby required), business-critical (warm standby appropriate), and operational (cold standby or backup-and-restore acceptable). The mistake is treating every system as mission-critical by default, which inflates DR spend without improving recovery outcomes for lower-tier systems.

Classify workloads by asking what a four-hour outage actually costs. A payment processing gateway that handles $2 million per hour in transactions clearly warrants hot standby. A monthly reporting pipeline does not. That four-hour cost calculation should sit next to your DR infrastructure cost before any architecture decision is approved.

This tiering approach also shapes how Australian businesses should structure their multi-region cloud strategy, since the data residency and latency constraints of a second Australian region behave differently depending on whether you're running hot or warm.

Configuration drift: the underrated DR risk

Neither warm nor hot standby provides reliable recovery if your standby environment has drifted from production. This is one of the most common failure modes in DR testing. The standby was correctly configured at deployment, then production was updated, patched, or reconfigured without the changes being propagated to the standby. When failover happens, the promoted environment doesn't behave like production and the incident extends.

Infrastructure-as-code is the primary mitigation. Terraform, Pulumi, and AWS CloudFormation all support multi-region deployment from a single codebase, which keeps both environments in sync if the pipeline is disciplined. Without that discipline, warm and hot standby both become warm fiction.

Testing frequency: what actually separates teams that survive from those that don't

DR models are only as reliable as the last successful test. Australian organisations that test quarterly or annually discover, under real failure conditions, that their standby assumptions were wrong. Failover scripts that worked 18 months ago don't account for new dependencies. Database replication jobs that were healthy have silently fallen behind.

Hot standby environments should be tested by actually routing production traffic to them, at least in a controlled canary configuration, every quarter. Warm standby environments should undergo full promotion tests at the same cadence. The test should include the full promotion sequence: DNS cutover, capacity scale-up, and database promotion. Each of those steps has its own failure modes, and finding them under controlled conditions is the entire point.

Document the test results against your RTO and RPO targets. If the warm standby consistently promotes in 12 minutes and your RTO is 20 minutes, you have an 8-minute buffer. If a dependency change pushes that to 22 minutes, you know before a real outage forces the discovery.

Costs side by side

As a rough guide for a mid-tier production workload running on AWS Sydney with a warm or hot standby in AWS Melbourne: a warm standby at 25% capacity adds approximately 25–35% to total infrastructure cost. A hot standby at 100% capacity roughly doubles it. Those figures shift depending on reserved instance coverage, data transfer volumes, and whether your RDS configuration uses Multi-AZ or cross-region read replicas.

The right answer for most Australian IT teams is a mixed model: hot standby for two or three genuinely mission-critical systems, warm standby for the broader business-critical tier, and backup-and-restore for everything else. That approach matches recovery investment to actual business impact rather than treating DR as a uniform cost centre.

→ The Confirmations · Daily newsletter

One email at 06:00 UTC. Six minutes. The only digest written for desks, not for retail.