
Testing Multi-AZ Failover Assumptions After the Yandex Cloud Sasovo Outage
Introduction: Testing Failover Assumptions You Have Never Exercised
Two reports, both published on 11 October 2026, describe the same event. ET Datacenters puts it at roughly 00:35 UTC; Ukraine's UNN at roughly 04:25 UTC. Both say a Yandex Cloud zone went down after a drone attack on a Yandex data center in Sasovo. That is the entire confirmed record I have to work with.
What follows is not sympathy, and it is not a verdict on Yandex. It is a test plan for the failover assumptions sitting in your own manifests, and the useful engineering response to the Sasovo outage is to re-test the ones you have almost certainly never exercised. My position up front: for most teams, "multi-AZ" is a placement policy in a manifest, not a measured capability. You asked the scheduler to spread replicas across three zones. You never verified that the surviving two can carry the load, that your control plane still answers, or that failover automation finishes inside the SLO you put in the contract.
Sasovo is a reminder of what topology diagrams hide. A zone is not an abstraction floating above geography. It terminates in buildings, power feeds, cooling plants, and network paths — and a physical event at one building can be a failure domain shared by the zones that depend on it.
What the Reports Actually Say — and What Is Unconfirmed
I want to be precise about the evidence here, because most incident commentary isn't. This post was built from two news report summaries, not from a full incident report, a provider postmortem, or a vendor advisory. Both summaries say a Yandex Cloud zone went down following an attack on the Sasovo data center. They agree on cause and facility. The ET Datacenters item lands about four hours before UNN's.
Everything beyond that is inference or unknown.
Open Questions About the Outage and What Would Resolve Them
| Unconfirmed item | Artifact that would settle it |
|---|---|
| Which zone identifier failed | A Yandex Cloud status-page postmortem naming the zone |
| Outage duration | Status-page timeline with start/recovery timestamps |
| Whether customer data was lost | Official statement or customer-reported recovery timelines |
| Whether automated failover triggered | Provider postmortem plus customer-side failover logs |
| Whether the failed asset was the facility, the network path into it, or power/cooling | Root-cause section of a postmortem, or independent reporting from the site |
That last row matters more than it looks. If the network path into the site failed rather than the site itself, remediation shifts from site redundancy to path diversity — different engineering, different cost, different questions for the vendor. The public reports don't draw that distinction, so I won't either.
Multi-AZ Is a Placement Claim, Not a Tested Capability
What a Zone Actually Promises
Providers document availability zones as separate fault domains inside a region: independent power, cooling, and network. That is a documented property of their infrastructure. It says nothing about your system. The promise does not cover the provider's global control plane, IAM, DNS, managed-service backends, or the cross-zone backbone.
And the backbone is shared by definition. Cross-zone replication is a feature that runs over the interconnect between zones. The link that gives you durability is itself a shared dependency, and it rarely appears in anyone's dependency register.
The Shared Failure Domains Cloud Abstractions Hide
Even with workloads spread cleanly across zones, these layers are commonly shared:
- the cloud API and control plane
- the identity provider and token issuance
- the managed database's replication path and its own control plane
- the container registry and the secret store
- DNS and the global load balancer
- the provider account itself — quotas, IAM, billing
- the region's physical exposure to a single event
Multi-AZ is a property of your data plane. Almost everything you use to operate that data plane is not zone-scoped.
Building an Assumption Inventory You Can Falsify
Convert vague resilience claims into falsifiable assertions. An assumption with no falsification test is a slogan.
| Assumption | Current evidence | Test that would falsify it |
|---|---|---|
| Losing one zone does not affect writes | Provider docs | Drain one zone in a non-prod account, watch write latency and errors |
| Failover completes inside our SLO | A runbook nobody has timed | Time a manual failover, separately from infrastructure timing |
| We have capacity to absorb a full zone | HPA max replicas, on paper | 1.5x load test against surviving zones only |
| Our control plane is not on the failed zone's path | None | Block API access from a test subnet, watch controllers |
| Backups are recoverable | Nightly job shows "success" | Restore into a different account and diff a row count |
From Documentation to Your Real Topology
Dump what you deployed, not what you intended. These commands are read-only:
# Kubernetes: where are pods actually scheduled?
kubectl get pods -A -o custom-columns='NODE:.spec.nodeName,NAMESPACE:.metadata.namespace,POD:.metadata.name' | head -20
## Node-to-zone mapping
kubectl get nodes -L topology.kubernetes.io/zone
## Cloud-side: instances, subnets, managed-service primaries per zone
## (Yandex Cloud CLI shown; subcommand names vary by provider and version,
## and every command below is read-only)
yc compute instance list
yc vpc subnet list
yc managed-postgresql cluster listThe output habitually reveals workloads pinned to one zone — by persistent-volume affinity, by on-demand instance availability that exists in only one zone, or by a hardcoded regional endpoint in a config file that predates the resilience plan.
Dependency Fan-Out You Forgot to Classify
Walk the call graph outward from the application and label each dependency zone-scoped, region-scoped, or global: secret manager, image registry, object storage, managed queue, managed database, identity, CI/CD, and any SaaS you reach over a single-region endpoint. Only the zone-scoped ones are covered by a multi-AZ claim. A region-scoped managed service does not become redundant because your compute is spread across three zones — that is the single most common category error I see in architecture reviews.
Testing Multi-AZ Failover for Real
Escalate. Read-only audits first, then control-plane degradation, then a single-zone drain in a non-production account, then scheduled production game days. Safety boundary: run these against accounts you own. Never against a third party's infrastructure, and never against a provider you are not paying.
Step 1 — Separate Control Plane From Data Plane
Degrade or remove API access to one zone while workloads keep running. The failure this catches is the one where the application is healthy but nothing can be scaled, restarted, or reconfigured — and where controllers that reconcile continuously begin evicting otherwise fine nodes because they cannot reach the API. Practical version: block the API endpoint from a test subnet with a firewall rule, or revoke a narrowly scoped token, and watch what the controllers do.
Step 2 — Drain a Single Zone Under Real Traffic
Cordon and drain one zone's nodes at a controlled rate while measuring error rate, latency, and in-flight connection failures.
ZONE="zone-a" # confirmed from: kubectl get nodes -L topology.kubernetes.io/zone
RATE_SECONDS=60
for node in $(kubectl get nodes -l topology.kubernetes.io/zone=$ZONE -o name); do
kubectl cordon "$node"
kubectl drain "$node" --ignore-daemonsets --delete-emptydir-data --timeout=300s
sleep $RATE_SECONDS # deliberate: do not drain the whole zone at once
doneMeasure with a query you already trust:
sum(rate(http_requests_total{code=~"5.."}[1m]))
/ sum(rate(http_requests_total[1m]))
Illustrative shape of a three-node zone drain when capacity is not actually N+1 — these numbers are a template for the output format, not measurements from the Sasovo event, which has no public customer-level data:
cordoned/drained zone-a 3/3 nodes
p99 latency 210ms -> 265ms
5xx rate 0.04% -> 0.31%
in-flight resets 0 -> 148
A 5xx step that does not recover when the drain completes is your answer.
Step 3 — Measure RTO and RPO Instead of Assuming Them
RTO is wall-clock time from failure to the service meeting its latency and error targets again. RPO is the window of writes you lost. Timestamp the last good write before failover, then the first good write after:
SELECT max(committed_at) AS last_good_before
FROM orders WHERE committed_at < :failover_start;
SELECT min(committed_at) AS first_good_after
FROM orders WHERE committed_at > :failover_start;
The gap is your measured RPO. Any number presented in this post is a method, not a measurement.
Step 4 — Capacity Headroom and Quota Math
If one of three zones disappears, survivors must absorb 1.5x load — plus retry amplification and reconnect storms, which routinely push the real multiplier past 1.5. Check instance quotas, IP address quotas, autoscaler maximums, database connection ceilings, and storage throughput limits in the surviving zones. Instance headroom is not IP headroom, and neither is database connection headroom. Those are separately capped and are frequently the binding constraint.
Step 5 — The Dependencies That Fail Silently
Some failure modes emit no error at all: long DNS TTLs keep clients pointed at a dead zone; connection pools fill with dead sockets; retries synchronize into a thundering herd; caches stampede after a cold start; leader election never completes because two-of-two voters became one. The mitigations are unglamorous — short TTLs with a fallback resolver, jittered exponential backoff, circuit breakers, pool recycling on idle timeouts — but they are the difference between degradation and an outage.
Region Isolation and the Limits of a Single Provider
Separate three scenarios: a zone failure, a whole-region failure, and a provider-level event. Multi-AZ does nothing for the second or third. The Sasovo reports are a reminder that a provider region has a physical location, and that physical locations are affected by events no architecture diagram represents.
Redundancy That Shares a Single Point of Failure
- Three zones behind one global load balancer and one DNS zone → check the DNS provider's blast radius.
- Three zones in one account with one credential set → check whether losing the account's IAM is survivable.
- "Multi-provider" where both providers sit behind the same CDN or transit path → check the traceroute.
- Backups stored inside the same provider they protect against → check where a restore would run.
Portability Tax and Data Gravity
Be honest about exit cost. Managed database engines, proprietary queues, IAM models, and egress fees decide whether a second provider is a real option or a slide. Assessment method: pick the two or three services with the highest switching cost, and estimate migration in engineer-weeks. Then compare that number to the probability you will need it. Both numbers are guesses; writing them down is still better than adjectives.
What You Cannot Test, and What to Do Instead
This section is judgment, not something derived from the reports. Physical attacks, facility loss, sanctions, and provider-wide incidents cannot be staged. The substitute is preparation you can rehearse.
Define a Degraded-Mode Contract
Write down what the product does with one zone gone: which features degrade, which are disabled, what users see, and who decides to enter the mode. This converts an untestable scenario into a testable one, because degraded mode can be exercised on demand, on a Tuesday afternoon, without any disaster.
Manual Failover Runbook and Game Day
Rehearse failover by hand and time it. Record human latency separately from infrastructure latency — the runbook is usually the slowest component. It needs an authoritative topology, credentials escrowed outside the failed domain, rollback steps, and one named decision owner.
Turning Test Results Into a Decision
Write the measured RTO and RPO next to the SLO. Then choose exactly one of: accept (and record that you accepted, with the number), remediate (with an owner and a date), or re-architect (with a budget). The default outcome — a test that produces a document and no change — is the most expensive result on the list, because it consumes the credibility you will need for the next one.
Conclusion
The Sasovo reports are a prompt to verify, not a verdict on one vendor. The assumption worth re-testing is the one you have never actually exercised, and it is probably the one in your manifest that says topologySpreadConstraints and gets no further thought. Drain one zone in a non-production environment this week. Write down the number you get.
Further Reading
- Yandex cloud zone down after Sasovo data center drone attack — ET Datacenters report summary, 11 October 2026, ~00:35 UTC.
- Russia reports a major Yandex Cloud outage following an attack on a Yandex data center — UNN report summary, 11 October 2026, ~04:25 UTC.
- AWS: Regions and Availability Zones — an example of the standard zone-isolation promise that providers document and that teams routinely over-read.


