Live · Wed, Oct 7, 2026 · 06:01 UTC Block 843,917 Fees 14 sat/vB Fear & Greed 72 · Greed
Newsletter Pro Terminal Sign in
ITop Field News.
Subscribe →
Live · 06:01 UTC Block 843,917 F&G 72
Cloud & infrastructure Cloud & infrastructure desk

Cloud availability zones: what they are and when they matter

Cloud availability zones are widely misunderstood as simple backup locations, but the design choices you make around them determine whether your workloads survive a real failure. Here's what Australian IT teams need to know.

Close-up of cooling fans in a server room, showcasing technology and efficiency.

Photo by panumas nikhomkhai on Pexels

Cloud availability zones are one of those infrastructure concepts that most teams assume they understand until an outage exposes the gap. On paper, the idea sounds simple: spread your workloads across multiple physical facilities within a region so that one failure doesn't take everything down. In practice, the decisions about which workloads go where, how data replicates between zones, and what trade-offs you're accepting in latency and cost are far more consequential than any vendor diagram suggests.

What availability zones actually are

An availability zone (AZ) is a physically separate data centre, or cluster of data centres, within a single cloud region. AWS, Azure, and GCP all offer AZs in their Australian regions. AWS Sydney offers three AZs. Azure's Australian East region (Sydney) offers three as well. Google Cloud's Sydney region similarly provides three zones. Each AZ runs its own independent power supply, cooling, and network connections. They're close enough to each other to keep inter-zone network latency low (typically under two milliseconds), but physically far enough apart that a localised event like a power grid failure or flooding won't simultaneously affect more than one.

The key word is "typically." Vendors don't publish the physical distance between their AZs, and the 2021 AWS ap-southeast-2 outage demonstrated that zone-level failures can still cascade in unexpected ways, particularly where shared control plane components are involved. An AZ is not an airtight blast radius. It reduces correlated failure risk. It doesn't eliminate it.

How AZs differ from regions

This is where many teams conflate two different things. A region is a geographic cluster of AZs. A zone is a single facility inside that cluster. Deploying across multiple AZs keeps you inside one region. Deploying across multiple regions means running in entirely separate geographic locations, like Sydney and Melbourne, or Sydney and Singapore.

Multi-AZ deployments protect against data centre-level failures. Multi-region deployments protect against regional outages, network partitions, or catastrophic events. The protection level is dramatically different, and so is the cost. Most workloads that don't have a genuine geographic requirement are better served by a solid multi-AZ architecture than an under-resourced multi-region one. Trying to be everywhere at once, with insufficient automation and runbook maturity to manage it, often produces worse resilience than a focused two-AZ design.

For a deeper look at how these decisions interact with disaster recovery commitments, the comparison between warm standby and hot standby DR models is worth reading alongside this piece, since AZ design directly determines which recovery architecture is even feasible.

Where Australian teams get it wrong

Three mistakes come up repeatedly in Australian enterprise deployments.

First, teams deploy to multiple AZs but forget to make their application stateless or to synchronise state correctly between zones. A web tier spread across three AZs is useless if session data lives on a single in-memory cache that doesn't replicate. The application layer needs to be designed for zone failure, not just the infrastructure underneath it.

Second, teams use AZ labels as a proxy for isolation, assuming that "AZ-a" and "AZ-b" in two different accounts are actually in different physical facilities. They're not guaranteed to be. AWS randomises AZ label assignments per account precisely to distribute load. If you need genuine physical separation between two tenants, you need to use AZ IDs (the stable identifiers like apse2-az1), not AZ names.

Third, teams over-provision for AZ failure at the infrastructure layer but underprepare their runbooks and automation. A multi-AZ database cluster is only useful in an outage if the failover is tested, fast, and doesn't require manual DNS changes that take 30 minutes to propagate.

Latency and cost: the real trade-offs

Inter-AZ traffic isn't free. AWS charges for data transferred between AZs inside the same region. On high-throughput workloads, particularly those with lots of internal service-to-service calls, inter-AZ data transfer can become a meaningful line item. This is a known cost driver that teams often miss during architecture reviews.

Latency across AZs is low but not zero. For most application patterns it's irrelevant. For synchronous distributed transactions, particularly those using consensus protocols like Raft or Paxos, the 1–2 ms round trip per AZ hop adds up across multiple coordination steps. This doesn't mean you shouldn't use multiple AZs; it means you should understand your workload's sensitivity to that latency before committing to a design that relies on tight cross-AZ coordination.

These hidden costs are part of a broader pattern worth examining: cloud egress and transfer costs in Australia catch many teams off guard precisely because they don't appear in headline pricing.

What a good multi-AZ architecture looks like

For a standard three-tier application running on AWS Sydney, a reasonable baseline looks like this:

  • Load balancer deployed across all three AZs, with cross-zone load balancing enabled.
  • Application instances in at least two AZs, with auto-scaling groups configured to maintain balance after a zone failure.
  • Managed database in Multi-AZ mode (RDS Multi-AZ, for example), with synchronous replication and automatic failover.
  • Caching layer (ElastiCache, Redis) deployed with cluster mode enabled and replicas spread across AZs.
  • Health checks that actually test application health, not just TCP connectivity.

The last point matters more than teams expect. A health check that passes because port 443 is open tells you nothing about whether the application can actually serve requests. Shallow health checks mean your load balancer keeps routing traffic to a zone that's technically up but functionally broken.

AZ design and the Essential Eight

The Australian Signals Directorate's Essential Eight doesn't directly mandate multi-AZ architecture, but the resilience it's meant to produce is directly relevant to the "regular backups" and "patch operating systems" controls. Systems that can't survive an AZ failure often can't survive a ransomware-triggered shutdown either. Availability design and security design are the same problem from different angles. Teams that treat them as separate workstreams end up with well-secured systems that still go down when a power transformer fails in Mascot.

Sovereign cloud requirements add another layer. Some Australian government and regulated-sector workloads are constrained to specific AZs or physical facilities by data residency requirements. In those cases, multi-AZ architecture may not span all three available zones; it may be limited to two, or even one, depending on how the sovereign cloud provider has built out its facilities. Understanding the physical geography of your AZs is not optional for these workloads.

When a single AZ is actually acceptable

Not every workload needs multi-AZ deployment. Development and test environments, batch jobs with no SLA, and non-critical internal tooling are all candidates for single-AZ deployment where cost discipline matters. The mistake isn't choosing a single AZ for low-criticality workloads; it's failing to make that choice explicitly. When an AZ is chosen by default rather than by design, teams often discover during an incident that the workload turned out to matter more than they thought.

Classify your workloads. Assign them a criticality tier. Map that tier to an AZ requirement. Document it. This sounds obvious, but most Australian organisations that have gone through a cloud migration have at least a handful of systems where nobody is quite sure what the recovery expectation is. That ambiguity is where outages become incidents and incidents become crises.

→ The Confirmations · Daily newsletter

One email at 06:00 UTC. Six minutes. The only digest written for desks, not for retail.