Testing Multi-AZ Failover Assumptions After the Yandex Cloud Sasovo Outage

Testing Multi-AZ Failover Assumptions After the Yandex Cloud Sasovo Outage

pr0h0•
cloud-resiliencedisaster-recoveryyandex-cloudmulti-azinfrastructure-risk
AI Usage (96%)

Introduction: Testing Failover Assumptions You Have Never Exercised

Two reports, both published on 11 October 2026, describe the same event. ET Datacenters puts it at roughly 00:35 UTC; Ukraine's UNN at roughly 04:25 UTC. Both say a Yandex Cloud zone went down after a drone attack on a Yandex data center in Sasovo. That is the entire confirmed record I have to work with.

What follows is not sympathy, and it is not a verdict on Yandex. It is a test plan for the failover assumptions sitting in your own manifests, and the useful engineering response to the Sasovo outage is to re-test the ones you have almost certainly never exercised. My position up front: for most teams, "multi-AZ" is a placement policy in a manifest, not a measured capability. You asked the scheduler to spread replicas across three zones. You never verified that the surviving two can carry the load, that your control plane still answers, or that failover automation finishes inside the SLO you put in the contract.

Sasovo is a reminder of what topology diagrams hide. A zone is not an abstraction floating above geography. It terminates in buildings, power feeds, cooling plants, and network paths — and a physical event at one building can be a failure domain shared by the zones that depend on it.

What the Reports Actually Say — and What Is Unconfirmed

I want to be precise about the evidence here, because most incident commentary isn't. This post was built from two news report summaries, not from a full incident report, a provider postmortem, or a vendor advisory. Both summaries say a Yandex Cloud zone went down following an attack on the Sasovo data center. They agree on cause and facility. The ET Datacenters item lands about four hours before UNN's.

Everything beyond that is inference or unknown.

Open Questions About the Outage and What Would Resolve Them

Unconfirmed itemArtifact that would settle it
Which zone identifier failedA Yandex Cloud status-page postmortem naming the zone
Outage durationStatus-page timeline with start/recovery timestamps
Whether customer data was lostOfficial statement or customer-reported recovery timelines
Whether automated failover triggeredProvider postmortem plus customer-side failover logs
Whether the failed asset was the facility, the network path into it, or power/coolingRoot-cause section of a postmortem, or independent reporting from the site

That last row matters more than it looks. If the network path into the site failed rather than the site itself, remediation shifts from site redundancy to path diversity — different engineering, different cost, different questions for the vendor. The public reports don't draw that distinction, so I won't either.

Multi-AZ Is a Placement Claim, Not a Tested Capability

What a Zone Actually Promises

Providers document availability zones as separate fault domains inside a region: independent power, cooling, and network. That is a documented property of their infrastructure. It says nothing about your system. The promise does not cover the provider's global control plane, IAM, DNS, managed-service backends, or the cross-zone backbone.

And the backbone is shared by definition. Cross-zone replication is a feature that runs over the interconnect between zones. The link that gives you durability is itself a shared dependency, and it rarely appears in anyone's dependency register.

The Shared Failure Domains Cloud Abstractions Hide

Even with workloads spread cleanly across zones, these layers are commonly shared:

  • the cloud API and control plane
  • the identity provider and token issuance
  • the managed database's replication path and its own control plane
  • the container registry and the secret store
  • DNS and the global load balancer
  • the provider account itself — quotas, IAM, billing
  • the region's physical exposure to a single event

Multi-AZ is a property of your data plane. Almost everything you use to operate that data plane is not zone-scoped.

Building an Assumption Inventory You Can Falsify

Convert vague resilience claims into falsifiable assertions. An assumption with no falsification test is a slogan.

AssumptionCurrent evidenceTest that would falsify it
Losing one zone does not affect writesProvider docsDrain one zone in a non-prod account, watch write latency and errors
Failover completes inside our SLOA runbook nobody has timedTime a manual failover, separately from infrastructure timing
We have capacity to absorb a full zoneHPA max replicas, on paper1.5x load test against surviving zones only
Our control plane is not on the failed zone's pathNoneBlock API access from a test subnet, watch controllers
Backups are recoverableNightly job shows "success"Restore into a different account and diff a row count

From Documentation to Your Real Topology

Dump what you deployed, not what you intended. These commands are read-only:

topology-dump.sh
# Kubernetes: where are pods actually scheduled?
kubectl get pods -A -o custom-columns='NODE:.spec.nodeName,NAMESPACE:.metadata.namespace,POD:.metadata.name' | head -20

## Node-to-zone mapping
kubectl get nodes -L topology.kubernetes.io/zone

## Cloud-side: instances, subnets, managed-service primaries per zone
## (Yandex Cloud CLI shown; subcommand names vary by provider and version,
## and every command below is read-only)
yc compute instance list
yc vpc subnet list
yc managed-postgresql cluster list

The output habitually reveals workloads pinned to one zone — by persistent-volume affinity, by on-demand instance availability that exists in only one zone, or by a hardcoded regional endpoint in a config file that predates the resilience plan.

Dependency Fan-Out You Forgot to Classify

Walk the call graph outward from the application and label each dependency zone-scoped, region-scoped, or global: secret manager, image registry, object storage, managed queue, managed database, identity, CI/CD, and any SaaS you reach over a single-region endpoint. Only the zone-scoped ones are covered by a multi-AZ claim. A region-scoped managed service does not become redundant because your compute is spread across three zones — that is the single most common category error I see in architecture reviews.

Testing Multi-AZ Failover for Real

Escalate. Read-only audits first, then control-plane degradation, then a single-zone drain in a non-production account, then scheduled production game days. Safety boundary: run these against accounts you own. Never against a third party's infrastructure, and never against a provider you are not paying.

Step 1 — Separate Control Plane From Data Plane

Degrade or remove API access to one zone while workloads keep running. The failure this catches is the one where the application is healthy but nothing can be scaled, restarted, or reconfigured — and where controllers that reconcile continuously begin evicting otherwise fine nodes because they cannot reach the API. Practical version: block the API endpoint from a test subnet with a firewall rule, or revoke a narrowly scoped token, and watch what the controllers do.

Step 2 — Drain a Single Zone Under Real Traffic

Cordon and drain one zone's nodes at a controlled rate while measuring error rate, latency, and in-flight connection failures.

drain-one-zone.sh
ZONE="zone-a"          # confirmed from: kubectl get nodes -L topology.kubernetes.io/zone
RATE_SECONDS=60

for node in $(kubectl get nodes -l topology.kubernetes.io/zone=$ZONE -o name); do
kubectl cordon "$node"
kubectl drain "$node"   --ignore-daemonsets   --delete-emptydir-data   --timeout=300s
sleep $RATE_SECONDS   # deliberate: do not drain the whole zone at once
done

Measure with a query you already trust:

sum(rate(http_requests_total{code=~"5.."}[1m]))
  / sum(rate(http_requests_total[1m]))

Illustrative shape of a three-node zone drain when capacity is not actually N+1 — these numbers are a template for the output format, not measurements from the Sasovo event, which has no public customer-level data:

cordoned/drained   zone-a 3/3 nodes
p99 latency        210ms -> 265ms
5xx rate           0.04% -> 0.31%
in-flight resets   0 -> 148

A 5xx step that does not recover when the drain completes is your answer.

Step 3 — Measure RTO and RPO Instead of Assuming Them

RTO is wall-clock time from failure to the service meeting its latency and error targets again. RPO is the window of writes you lost. Timestamp the last good write before failover, then the first good write after:

SELECT max(committed_at) AS last_good_before
FROM orders WHERE committed_at < :failover_start;

SELECT min(committed_at) AS first_good_after
FROM orders WHERE committed_at > :failover_start;

The gap is your measured RPO. Any number presented in this post is a method, not a measurement.

Step 4 — Capacity Headroom and Quota Math

If one of three zones disappears, survivors must absorb 1.5x load — plus retry amplification and reconnect storms, which routinely push the real multiplier past 1.5. Check instance quotas, IP address quotas, autoscaler maximums, database connection ceilings, and storage throughput limits in the surviving zones. Instance headroom is not IP headroom, and neither is database connection headroom. Those are separately capped and are frequently the binding constraint.

Step 5 — The Dependencies That Fail Silently

Some failure modes emit no error at all: long DNS TTLs keep clients pointed at a dead zone; connection pools fill with dead sockets; retries synchronize into a thundering herd; caches stampede after a cold start; leader election never completes because two-of-two voters became one. The mitigations are unglamorous — short TTLs with a fallback resolver, jittered exponential backoff, circuit breakers, pool recycling on idle timeouts — but they are the difference between degradation and an outage.

Region Isolation and the Limits of a Single Provider

Separate three scenarios: a zone failure, a whole-region failure, and a provider-level event. Multi-AZ does nothing for the second or third. The Sasovo reports are a reminder that a provider region has a physical location, and that physical locations are affected by events no architecture diagram represents.

Redundancy That Shares a Single Point of Failure

  • Three zones behind one global load balancer and one DNS zone → check the DNS provider's blast radius.
  • Three zones in one account with one credential set → check whether losing the account's IAM is survivable.
  • "Multi-provider" where both providers sit behind the same CDN or transit path → check the traceroute.
  • Backups stored inside the same provider they protect against → check where a restore would run.

Portability Tax and Data Gravity

Be honest about exit cost. Managed database engines, proprietary queues, IAM models, and egress fees decide whether a second provider is a real option or a slide. Assessment method: pick the two or three services with the highest switching cost, and estimate migration in engineer-weeks. Then compare that number to the probability you will need it. Both numbers are guesses; writing them down is still better than adjectives.

What You Cannot Test, and What to Do Instead

This section is judgment, not something derived from the reports. Physical attacks, facility loss, sanctions, and provider-wide incidents cannot be staged. The substitute is preparation you can rehearse.

Define a Degraded-Mode Contract

Write down what the product does with one zone gone: which features degrade, which are disabled, what users see, and who decides to enter the mode. This converts an untestable scenario into a testable one, because degraded mode can be exercised on demand, on a Tuesday afternoon, without any disaster.

Manual Failover Runbook and Game Day

Rehearse failover by hand and time it. Record human latency separately from infrastructure latency — the runbook is usually the slowest component. It needs an authoritative topology, credentials escrowed outside the failed domain, rollback steps, and one named decision owner.

Turning Test Results Into a Decision

Write the measured RTO and RPO next to the SLO. Then choose exactly one of: accept (and record that you accepted, with the number), remediate (with an owner and a date), or re-architect (with a budget). The default outcome — a test that produces a document and no change — is the most expensive result on the list, because it consumes the credibility you will need for the next one.

Conclusion

The Sasovo reports are a prompt to verify, not a verdict on one vendor. The assumption worth re-testing is the one you have never actually exercised, and it is probably the one in your manifest that says topologySpreadConstraints and gets no further thought. Drain one zone in a non-production environment this week. Write down the number you get.

Further Reading

Share this post

More posts

Comments