
Testing Active-Active and Warm-Standby Assumptions with an Azure Blast-Radius Drill
Introduction: What the 40-Hour, 18-Region Report Actually Challenges
This post is a hands-on blast-radius drill for testing two multi-region assumptions that ride along with almost every design and almost never get falsified: that active-active makes a single-region loss invisible, and that warm standby promotes in minutes. Both are cheap to type into a design doc. Both are expensive to prove wrong.
That is the point of the drill below — to falsify them against a correlated failure pattern. The shape comes from the shattered.io write-up published on 2026-10-02: two Azure incidents roughly 40 hours apart, with the affected set reported at 18 regions. If that reporting holds up, plenty of teams with "multi-region" stamped on their architecture had multiple regions fail inside the same recovery window — the exact case their failover plan assumes away.
This is not an attempt to reproduce a platform outage. It is a scoped blast-radius drill you run in your own non-production estate to find out which recovery assumptions were never true: control-plane reachability, token issuance, DNS and traffic management, registry pulls, quota headroom, and whether your failback path has ever run in the direction you claim.
What the Azure Report Says, and What It Leaves Unanswered
Confirmed by the Source Report
Three things are stated plainly in the source material:
- Two Azure incidents occurred within roughly a 40-hour window.
- The affected scope is reported as 18 regions. The write-up is characterized as a report, not an Azure RCA, so "18 regions" is a claim I am preserving, not a fact I have verified.
- The write-up was published by shattered.io on 2026-10-02.
That is the whole evidentiary base. Everything else here is either engineering reasoning or something you can test on your own subscription.
Unknown Until a Primary Azure RCA Lands
The report does not give per-incident timelines, which services degraded in each incident, whether the blast radius came from the control plane, the data plane, or a shared singleton like identity or DNS, or whether regional isolation actually failed versus was never claimed for those services in the first place. As of this writing I have not seen a public Azure root cause analysis covering the window.
Those gaps matter more than the region count. A data-plane failure and a control-plane failure produce completely different symptoms in your application, and a shared-singleton failure produces a third set entirely. Until the RCA lands, the causality in every postmortem you read — including this one — is inference.
Blast Radius Is a Dependency Graph, Not a Region Count
Control Plane vs Data Plane Failure Modes
A resource can be healthy in three regions and unmanageable in all of them at once. That is the control-plane failure mode: the data plane keeps serving reads and writes from existing endpoints while the ARM/management plane — scaling, key rotation, role assignment, DNS record updates, new deployments — is unreachable. Monitoring says green. The runbook says red.
The reverse hurts too. A control plane that works fine plus a data plane with cross-region replication lag gives you a successful promotion into a region holding stale state.
The Shared Singletons That Collapse Regional Isolation
Regional isolation is a property of your dependency graph, not of the region boundary. The usual culprits:
| Shared dependency | Failure shape | What breaks even in a "healthy" region |
|---|---|---|
| Entra ID token issuance | global/regional control path | New managed-identity tokens, cold-start auth, workload identity federation |
| DNS, Traffic Manager, Front Door | shared control of name resolution | Traffic shift, endpoint discovery, private endpoint resolution |
| Container registry pulls | often single-region or single-instance | Scale-out cannot pull images, so failover capacity never starts |
| Key Vault | regional resource, cross-region references | Secret and certificate fetch at cold start |
| Regional quota and SKU capacity | per-subscription, per-region | Scale-out requests denied precisely when you need them |
| Status and health plane | shared | You cannot see the outage you are trying to respond to |
That last row is the underrated one. If your incident detection depends on a health endpoint sitting in the same failing plane, your mean time to detect is the length of the outage, not the length of your probe interval.
Designing the Drill: Probe, Blast Unit, and Blast Radius
Step 1 — Inventory Cross-Region Dependencies with Azure Resource Graph
Tag first, query second. A resource with no failover-role tag is invisible to the query below, which means it is invisible to your runbook too.
az extension add --name resource-graph
az graph query -q "
Resources
| extend role = tostring(tags['failover-role'])
| where isnotempty(role)
| project name, type, location, role, resourceGroup, subscriptionId
| order by role asc, location asc
" --output table
Illustrative output shape (trimmed, from a small lab subscription — yours will be longer):
| name | type | location | role |
|---|---|---|---|
| pg-orders-neu | microsoft.dbforpostgresql/flexibleservers | germanywestcentral | primary |
| pg-orders-weu | microsoft.dbforpostgresql/flexibleservers | westeurope | standby |
| acr-platform | microsoft.containerregistry/registries | westeurope | shared |
| kv-app-weu | microsoft.keyvault/vaults | westeurope | shared |
Two caveats worth stating: Resource Graph only sees control-plane-visible resources, so anything you create dynamically at runtime will not appear here, and the healthresources table (on subscriptions with Service Health access) is worth querying separately for the incident window so you can line your own blast radius up against the platform's event list.
Step 2 — Deploy a Synthetic Probe That Fails Loudly
A ping against /healthz exercises none of the failover path. The probe should touch identity, config, data read/write, and one downstream call, deployed per region:
const BUDGET = { identity: 1500, config: 400, secret: 600, write: 800, read: 500, downstream: 1000 };
const hops = {
identity: async () => {
const r = await fetch(
"http://169.254.169.254/metadata/identity/oauth2/token" +
"?api-version=2018-02-01&resource=https%3A%2F%2Fvault.azure.net",
{ headers: { Metadata: "true" }, signal: AbortSignal.timeout(BUDGET.identity) }
);
if (!r.ok) throw new Error("imds " + r.status);
return (await r.json()).access_token;
},
config: async () => {
const r = await fetch(process.env.CONFIG_URL, {
signal: AbortSignal.timeout(BUDGET.config),
});
if (!r.ok) throw new Error("config " + r.status);
return r.json();
},
secret: async (token) => {
const r = await fetch(process.env.KV_SECRET_URL, {
headers: { Authorization: "Bearer " + token },
signal: AbortSignal.timeout(BUDGET.secret),
});
if (!r.ok) throw new Error("kv " + r.status);
return r.json();
},
};
async function timed(name, fn, ...args) {
const t0 = performance.now();
try {
await fn(...args);
return { hop: name, ok: true, ms: Math.round(performance.now() - t0) };
} catch (err) {
return { hop: name, ok: false, ms: Math.round(performance.now() - t0), error: err.message };
}
}
const run = async () => {
const token = await hops.identity().catch(() => null);
const results = [
await timed("config", hops.config),
token ? await timed("secret", hops.secret, token) : { hop: "secret", ok: false, error: "no-token" },
await timed("data-write", writeProbeRow),
await timed("data-read", readProbeRow),
await timed("downstream", callDownstream),
];
const verdict = results.every((r) => r.ok) ? "pass" : "fail";
console.log(JSON.stringify({ region: process.env.REGION, verdict, results }));
};The failing run below has the identity hop as the first casualty — that is the pattern to look for, a region-local component failing because of a region-external dependency:
{"region":"westeurope","verdict":"fail","results":[
{"hop":"config","ok":true,"ms":41},
{"hop":"secret","ok":false,"ms":1502,"error":"no-token"},
{"hop":"data-write","ok":true,"ms":73},
{"hop":"data-read","ok":true,"ms":46},
{"hop":"downstream","ok":false,"ms":1014,"error":"ETIMEDOUT"}]}
Step 3 — Inject the Fault
Three blast units, ascending risk:
- DNS-level. Point an internal CNAME at a black-hole record in a non-production subscription and watch how long the old answer persists. Measure the TTL you actually have, not the one you configured:
dig +noall +answer api.example.com. - Endpoint-level. Block a single dependency (registry, Key Vault, an egress FQDN) via firewall or network policy. Chaos Studio is the managed option here, and for AKS targets it exposes Chaos Mesh-derived faults including DNS and network faults. Check the current fault list in the docs before planning around a specific one — the set evolves.
- Quota-level. Do not try to exhaust real quota. Cap the standby region deliberately and try to scale into it, so you learn whether the ceiling exists before an incident finds out for you.
Define rollback before execution: the exact CLI command, the resource, the expected post-state, and a time limit after which the drill aborts automatically. A fault injector without a rollback contract is just an outage with extra steps.
Step 4 — Measure Failover and Failback as Separate Numbers
Failover and failback are different systems. Report them separately:
- RTO per direction, measured from fault injection to probe pass, not from human detection.
- Error-budget burn, as a percentage of the monthly budget the drill itself consumed.
- Duplicate writes, counted by giving the probe a deterministic idempotency key and counting rows with the same key afterward.
- Reconnect storm behaviour, tracked as new connection attempts and TLS handshakes per second during promotion.
Where Warm Standby Usually Breaks
Regional Quota and SKU Capacity Are Not Reserved
Warm standby is a configuration, not a reservation. The standby region's ceiling is a guess until you scale into it:
az vm list-usage --location westeurope --output table
If reported CurrentValue plus your planned scale-out exceeds Limit, your promotion is a throttling event. Capacity floors belong in the standby region as a monitored property, not as an assumption about the cloud.
Identity, Key Vault, and RBAC Propagation Lag
Three latency sources get bundled into one failover number, and separating them is the point:
- Managed identity token caching. A cached token can mask an identity-plane outage entirely until it expires, then fail all at once.
- Private endpoint DNS. A private endpoint record living in a zone the standby region does not resolve against fails closed, and it fails at cold start.
- RBAC propagation. A role assignment written during failover may not apply immediately, turning a failover into a permission error under load.
Cold Caches, Connection Pools, and Replication Lag
A standby region holds no warm connection pool, no in-process cache, no prepared statements. Promotion therefore turns a fast healthy region into a slow one: every request pays to rebuild state. If the primary also served writes, check replication lag before promoting, or you will accept a promotion that silently loses the last seconds of writes.
Reading the Results: A Pass/Fail Table for Failover Claims
Pass/Fail Table Layout
| Claim | Observed | Verdict |
|---|---|---|
| Active-active: traffic shift is invisible to users | Error rate spike during shift, no user-facing 5xx | Pass or fail on evidence, not intent |
| Warm standby: promotion completes in target RTO | Promotion required manual DNS, quota, and identity steps | Fail if any step was manual |
| Degraded mode: read-only operation survives write-plane loss | Writes failed closed, reads continued | Pass only if tested explicitly |
| Status-plane visibility: detection is independent of the failing plane | Detection depended on the affected region's probe | Fail |
Fill "Observed" from probe logs and the platform's own event data — not from memory written during the incident.
Separating Localized Failure from Retry Amplification
This is the difference between a regional dependency and self-inflicted global damage. If 40 clients each retry three times without jitter against a dependency that is down, you have turned one region's failure into 120 requests per interval, spread across every region sharing that dependency. The timing tells you which: genuine regional failures show a clean step at injection, retry amplification shows a rising sawtooth correlated with your retry configuration. I suspect a meaningful share of "18 regions affected" reporting in any multi-region incident includes this amplification effect. That is inference, and only platform telemetry can settle it.
Defense: Degraded Mode Beats Perfect Failover
Concrete Mitigations for Multi-Region Failover
- Bound every retry. Three attempts, exponential backoff with full jitter, and an explicit non-retryable set that excludes 4xx except 429.
- Break circuits per dependency, not per service. One circuit per Key Vault, registry, or identity endpoint, so a failing dependency stops draining your connection pool.
const withRetry = (fn, { attempts = 3, base = 100, cap = 2000 } = {}) => async (...args) => {
let lastErr;
for (let i = 0; i < attempts; i++) {
try {
return await fn(...args);
} catch (err) {
if (!isRetryable(err)) throw err;
lastErr = err;
await new Promise((r) => setTimeout(r, Math.random() * Math.min(cap, base * 2 ** i)));
}
}
throw lastErr;
};
- Count DNS TTLs as failover latency. Your effective RTO includes resolver caches you do not control; a 300-second TTL is a five-minute floor on any traffic-shift strategy.
- Cache credentials and ship static config so a control-plane outage does not stop a running fleet from starting new instances.
- Pre-provision capacity floors in the standby region and alert when actual capacity drops below them.
Run the drill's failback direction first. Teams rehearse failover and assume failback, then discover during the real event that failback is the direction nobody scripted.
Limits of This Drill and What Would Confirm the Report
Honest Boundaries of a Non-Production Drill
A non-production drill cannot reproduce a platform-level fault. You cannot simulate Entra ID token issuance being degraded for a region, or a traffic-management plane failing the way a real incident would. The 18-region correlation in the shattered.io write-up is reported, not verified here, and no drill in your subscription changes that.
What would change it: an Azure RCA covering the window. Four answers to watch for. First, whether each incident originated in a control-plane component, a data plane, or a shared singleton. Second, whether the second incident was a regression introduced by the first mitigation. Third, whether regional isolation was ever claimed for the affected services, or whether the assumption came entirely from customer-side design docs. Fourth, whether the region count reflects independent regional failures or client retry behaviour spreading one failure outward. Until those are published, treat every architectural conclusion drawn from the event — this post included — as a hypothesis with a test attached.
Further Reading
- Azure status — the primary place to date the incident window; the history view on that page carries the platform's own event records.
- Azure Outage 2026: 2 Incidents Hit 18 Regions in 40 Hours — the shattered.io write-up published 2026-10-02, the source for the 40-hour and 18-region claims in this post.
- Azure Resource Graph documentation — query language, tables, and limits for the Step 1 inventory.
- Azure Chaos Studio documentation — current fault library, target requirements, and safety constraints.
Vendor RCAs supersede press coverage. When Azure publishes its postmortem for the window, that document, not the reporting summarized here, becomes the reference.
Conclusion: Test the Failure Pattern, Not the Region Count
The number that matters is not 18. It is whether your system survives the pattern — correlated failures close enough together that your recovery window overlaps itself. The drill above answers four questions cheaply: which dependencies cross your region boundary, which of them your probe actually exercises, whether your standby region can scale when nothing else works, and whether failback was ever scripted by anyone. Run it in a non-production subscription with rollback defined first, and expect the useful finding to be uncomfortable: the assumption you were most confident about is usually the one that has never been executed.


