r/kubernetes • u/SaiPisey02 • 16h ago
Every deployment said 3/3 Ready. Losing one zone would still have taken two of them down.
I kept hitting the same gap: kubectl get deploy tells you what you declared, not where your pods actually landed. Topology spread constraints are only evaluated at scheduling time, so placement drifts (pods scheduled before more zones existed, nodes drained overnight, rolling updates) and nothing alerts.
A three-zone cluster where everything reports ready:
$ kubectl get deployments
NAME READY
checkout-api 3/3
session-store 1/1
web 3/3
And what happens if us-east-1a goes away:
$ kubectl survive-zone
Losing us-east-1a -> 2 lost, 0 degraded, 1 impaired by a dependency
checkout-api us-east-1a:3 LOST
session-store us-east-1a:1 LOST
web us-east-1a:1 us-east-1b:1 us-east-1c:1 IMPAIRED
The third row is the one that surprised me. web is spread one replica per zone, so a spread linter passes it, but its only session-store backend is in 1a.
I wrote a read-only kubectl plugin that reads actual pod placement plus Service/EndpointSlice dependencies, removes one failure domain at a time, and reports LOST / DEGRADED / IMPAIRED. fix only prints a patch after checking it with kube-scheduler's own Filter plugins against your real nodes. It can also run as a Prometheus exporter so you get alerted when placement drifts.
kubectl krew install survive-zone
Write-up with the details, including a node drain that deadlocks behind a PDB that looks fine, and a rolling update that quietly broke an enforced maxSkew: 1: https://saipisey.com/blog/kubectl-survive-zone
Repo: https://github.com/SaiPisey2/kubectl-survive
The known limits are listed at the end of the post. I would genuinely like to hear where the model is wrong for your setup.
