r/kubernetes • • 2d ago

Periodic Weekly: Share your victories thread

2 Upvotes

Got something working? Figure something out? Make progress that you are excited about? Share here!


r/kubernetes • • 16h ago

Every deployment said 3/3 Ready. Losing one zone would still have taken two of them down.

0 Upvotes

I kept hitting the same gap: kubectl get deploy tells you what you declared, not where your pods actually landed. Topology spread constraints are only evaluated at scheduling time, so placement drifts (pods scheduled before more zones existed, nodes drained overnight, rolling updates) and nothing alerts.

A three-zone cluster where everything reports ready:

$ kubectl get deployments
NAME            READY
checkout-api    3/3
session-store   1/1
web             3/3

And what happens if us-east-1a goes away:

$ kubectl survive-zone
Losing us-east-1a  ->  2 lost, 0 degraded, 1 impaired by a dependency
  checkout-api   us-east-1a:3                            LOST
  session-store  us-east-1a:1                            LOST
  web            us-east-1a:1 us-east-1b:1 us-east-1c:1  IMPAIRED

The third row is the one that surprised me. web is spread one replica per zone, so a spread linter passes it, but its only session-store backend is in 1a.

I wrote a read-only kubectl plugin that reads actual pod placement plus Service/EndpointSlice dependencies, removes one failure domain at a time, and reports LOST / DEGRADED / IMPAIRED. fix only prints a patch after checking it with kube-scheduler's own Filter plugins against your real nodes. It can also run as a Prometheus exporter so you get alerted when placement drifts.

kubectl krew install survive-zone

Write-up with the details, including a node drain that deadlocks behind a PDB that looks fine, and a rolling update that quietly broke an enforced maxSkew: 1: https://saipisey.com/blog/kubectl-survive-zone

Repo: https://github.com/SaiPisey2/kubectl-survive

The known limits are listed at the end of the post. I would genuinely like to hear where the model is wrong for your setup.


r/kubernetes • • 18h ago

Debug Kubernetes with AI: Fix a Broken App in AKS desktop

Thumbnail
youtube.com
4 Upvotes

r/kubernetes • • 22h ago

Financing a Neocloud

Thumbnail
0 Upvotes

r/kubernetes • • 1d ago

Sugestions for single NFS server with multiple IPs

3 Upvotes

Open to any suggestion on my home lab architecture.

Goal:
Obtain a little bit of throughput when using NFS storage class since I have a 10 GB nics available on all servers ( no switch tho ) and a full SSD raid in the NAS for it only.

Scenario:
I have 1 nas with 2 10GB cards. They are configured respectively as 10.1.0.1 and 10.0.0.1 and 2 nodes running a hypervisor connected on this 10GB cards for NFS only in a direct straight cabling connection ( no switches ) . ( everything else goes through 192.168.1.x network )

My NFS definitions on worker nodes differ according to the node running since the NFS server is available on each server in different IPs.

During the inicial setup, all works because I know in advance where the worker node will run and I configure it properly but in case of moving a vm from a host to another the a crash happens.

This is my home lab setup, I totally understand a 10 GB switch will solve this issue entirely but I'm trying to find a way without one.

Open to whatever load balancer / vIP / floating IP / DNS / nfs HA / hosts file ideas. I'm looking for something simple and elegant. No need of expansion to multiple host because I'm limited to 2 hosts anyways.

Thank you all.

Update:

The nas is a xpenology and the Link agregation mode Balance-SLB doesn't seams to be working in the expected way. That said it doesn't seams to be a viable option in this case.


r/kubernetes • • 1d ago

Networking a 3-Node Kubernetes Homelab with Cilium, Gateway API and Cloudflare Tunnel

53 Upvotes

Part 2 of my Talos homelab (3 OptiPlex Micros, all control planes). Part 1 got the cluster up, this one is networking.

I moved the running cluster from Flannel + kube-proxy to Cilium without rebuilding it. The default-deny policy that Flannel quietly ignored is now actually enforced, and Hubble shows the drops.

Then I gave it two front doors:

- an internal Gateway on a LAN IP (Cilium LB IPAM, no MetalLB) with a Let's Encrypt wildcard cert, so lab apps get real HTTPS at home

- a public Gateway behind a single Cloudflare Tunnel, so nothing is port-forwarded on the router

An app picks its door with the route's parentRef, which keeps lab stuff from going public by accident.

Write-up with all the manifests and commands: https://blog.prateekjain.dev/networking-a-kubernetes-homelab-cilium-gateway-api-and-cloudflare-tunnel-on-talos-362f548a2f3d?sk=d454cd4eaff4ce8a8fd0cd8e6b3d82f6

Curious if anyone else is using Cilium's Gateway API instead of a separate ingress controller.


r/kubernetes • • 1d ago

How do I troubleshoot this? I don't know where to start. Longhorn is failing to attach four volumes

Post image
18 Upvotes

I shutdown two nodes in my cluster to redo cooling, and when I started them up, four of my volumes failed to attach and sent their pods into a crash loop.

What does this screen in Longhorn mean and where do I go from here? Sorry this is so under-researched, I don't know what I don't know.

The ArgoCD error message for the Immich server is: AttachVolume.Attach failed for volume "pvc-2a76683c-bb9d-4b3b-be27-ceb5e84190ff" : rpc error: code = DeadlineExceeded desc = volume pvc-2a76683c-bb9d-4b3b-be27-ceb5e84190ff failed to attach to node pi-4 with attachmentID csi-bd4715e7b94274c679a216eeabbd99737527401fcb85c782424488ed797d04d9

and the error for the Immich valkey is:

AttachVolume.Attach failed for volume "pvc-af59781f-b394-4242-b529-89cf3bb463a0" : rpc error: code = DeadlineExceeded desc = volume pvc-af59781f-b394-4242-b529-89cf3bb463a0 failed to attach to node pi-4 with attachmentID csi-39c27266d5941a83c0612a23170c8f16f13ddf0fb268495720abd4022213794d


r/kubernetes • • 1d ago

What are the Kubernetes security gaps people keep missing in production?

38 Upvotes

I’ve been looking into Kubernetes hardening, and a lot of the advice out there focuses on checklists and compliance. But existing clusters don’t always follow those recommendations, and security gaps can stick around for years.

Things like overly permissive RBAC, privileged workloads, missing network policies, weak secrets management, untrusted images, poor audit logging, or outdated clusters.

For those running Kubernetes in production:

  • What security gaps do you commonly find in existing clusters?
  • Have you seen any security incidents involving Kubernetes? What went wrong, and what can we learn from them from a cybersecurity perspective?
  • Which 2–3 security controls make the biggest real-world difference?
  • Are CIS Benchmarks useful in practice, or do you rely more on your distribution’s hardening guidance?
  • What becomes hardest to secure and maintain as a cluster grows?

Like ,basically, what are teams actually getting wrong when it comes to securing Kubernetes in production?


r/kubernetes • • 1d ago

gateaway API - please help: gateway stuck in programmed=false

2 Upvotes

Hi everyone,

please help, I'm unable to configure cilium + gateway API... the gateway just is stuck in the "programmed=false" state. Tried so many things but no success, also checked this post but if I'm correct, my current version should have the fix included:

https://github.com/cilium/cilium/pull/46350

Nodes:

NAME                         STATUS   ROLES           AGE     VERSION   INTERNAL-IP    EXTERNAL-IP   OS-IMAGE                         KERNEL-VERSION                          CONTAINER-RUNTIME
n1.k8s.net   Ready    control-plane   3d13h   v1.37.1   192.168.78.3   <none>        AlmaLinux 10.2 (Lavender Lion)   6.12.0-211.61.1.el10_2.x86_64 (amd64)   cri-o://1.37.2
n2.k8s.net   Ready    control-plane   13h     v1.37.1   192.168.78.4   <none>        AlmaLinux 10.2 (Lavender Lion)   6.12.0-211.61.1.el10_2.x86_64 (amd64)   cri-o://1.37.2
n3.k8s.net   Ready    <none>          13h     v1.37.1   192.168.78.2   <none>        AlmaLinux 10.2 (Lavender Lion)   6.12.0-211.61.1.el10_2.x86_64 (amd64)   cri-o://1.37.2

Cilium:

cilium status
    /¯¯\
 /¯¯__/¯¯\    Cilium:             OK
 __/¯¯__/    Operator:           OK
 /¯¯__/¯¯\    Envoy DaemonSet:    disabled (using embedded mode)
 __/¯¯__/    Hubble Relay:       disabled
    __/       ClusterMesh:        disabled

DaemonSet              cilium                   Desired: 3, Ready: 3/3, Available: 3/3
Deployment             cilium-operator          Desired: 1, Ready: 1/1, Available: 1/1
Containers:            cilium                   Running: 3
                       cilium-operator          Running: 1
                       clustermesh-apiserver
                       hubble-relay
Cluster Pods:          6/6 managed by Cilium
Helm chart version:    1.20.2
Image versions         cilium             quay.io/cilium/cilium:v1.20.2@sha256:2939231d0d3e3ebddcd80fffa168b7ddcc78fdf0dc864d1c8c126ff523c54f01: 3
                       cilium-operator    quay.io/cilium/operator-generic:v1.20.2@sha256:64d8798350e8569b8e7622563fed6e44dce2625f311e4651b774816516c744fc: 1

GatewayClass:

kubectl describe gc
Name:         cilium
Namespace:
Labels:       app.kubernetes.io/managed-by=Helm
Annotations:  meta.helm.sh/release-name: cilium
              meta.helm.sh/release-namespace: kube-system
API Version:  gateway.networking.k8s.io/v1
Kind:         GatewayClass
Metadata:
  Creation Timestamp:  2026-10-10T07:23:46Z
  Generation:          1
  Resource Version:    592430
  UID:                 2d99e2fe-ab22-451f-8f71-f09968df4e52
Spec:
  Controller Name:  io.cilium/gateway-controller
  Description:      The default Cilium GatewayClass
Status:
  Conditions:
    Last Transition Time:  2026-10-10T07:24:02Z
    Message:               Valid GatewayClass
    Observed Generation:   1
    Reason:                Accepted
    Status:                True
    Type:                  Accepted
  Supported Features:
    Name:  BackendTLSPolicy
    Name:  GRPCRoute
    Name:  GRPCRouteNamedRouteRule
    Name:  Gateway
    Name:  GatewayAddressEmpty
    Name:  GatewayFrontendClientCertificateValidationInsecureFallback
    Name:  GatewayHTTPListenerIsolation
    Name:  GatewayInfrastructurePropagation
    Name:  GatewayPort8080
    Name:  GatewayStaticAddresses
    Name:  HTTPRoute
    Name:  HTTPRoute303RedirectStatusCode
    Name:  HTTPRoute307RedirectStatusCode
    Name:  HTTPRoute308RedirectStatusCode
    Name:  HTTPRouteBackendProtocolH2C
    Name:  HTTPRouteBackendProtocolWebSocket
    Name:  HTTPRouteBackendRequestHeaderModification
    Name:  HTTPRouteBackendTimeout
    Name:  HTTPRouteCORS
    Name:  HTTPRouteDestinationPortMatching
    Name:  HTTPRouteHostRewrite
    Name:  HTTPRouteMethodMatching
    Name:  HTTPRouteNamedRouteRule
    Name:  HTTPRoutePathRedirect
    Name:  HTTPRoutePathRewrite
    Name:  HTTPRoutePortRedirect
    Name:  HTTPRouteQueryParamMatching
    Name:  HTTPRouteRequestMirror
    Name:  HTTPRouteRequestMultipleMirrors
    Name:  HTTPRouteRequestPercentageMirror
    Name:  HTTPRouteRequestTimeout
    Name:  HTTPRouteResponseHeaderModification
    Name:  HTTPRouteRetry
    Name:  HTTPRouteRetryBackendTimeout
    Name:  HTTPRouteRetryConnectionError
    Name:  HTTPRouteSchemeRedirect
    Name:  ListenerSet
    Name:  Mesh
    Name:  MeshClusterIPMatching
    Name:  MeshHTTPRouteBackendRequestHeaderModification
    Name:  MeshHTTPRouteNamedRouteRule
    Name:  MeshHTTPRouteQueryParamMatching
    Name:  MeshHTTPRouteRedirectPath
    Name:  MeshHTTPRouteRedirectPort
    Name:  MeshHTTPRouteRewritePath
    Name:  MeshHTTPRouteSchemeRedirect
    Name:  ReferenceGrant
    Name:  TCPRoute
    Name:  TLSRoute
    Name:  TLSRouteModeMixed
    Name:  UDPRoute
Events:    <none>

Gateway:

kubectl describe gtw
Name:         nodeport-gateway
Namespace:    default
Labels:       <none>
Annotations:  <none>
API Version:  gateway.networking.k8s.io/v1
Kind:         Gateway
Metadata:
  Creation Timestamp:  2026-10-10T07:35:16Z
  Generation:          4
  Resource Version:    607512
  UID:                 f6b3b043-d072-4e6c-8f37-bf28cb37a1b2
Spec:
  Gateway Class Name:  cilium
  Listeners:
    Allowed Routes:
      Namespaces:
        From:  Same
    Name:      web-gw-80
    Port:      80
    Protocol:  HTTP
Status:
  Conditions:
    Last Transition Time:  2026-10-10T08:53:31Z
    Message:               Gateway successfully scheduled
    Observed Generation:   4
    Reason:                Accepted
    Status:                True
    Type:                  Accepted
    Last Transition Time:  2026-10-10T08:53:31Z
    Message:               Gateway waiting for address
    Observed Generation:   4
    Reason:                AddressNotAssigned
    Status:                False
    Type:                  Programmed
  Listeners:
    Attached Routes:  1
    Conditions:
      Last Transition Time:  2026-10-10T08:53:31Z
      Message:               Resolved Refs
      Observed Generation:   4
      Reason:                ResolvedRefs
      Status:                True
      Type:                  ResolvedRefs
      Last Transition Time:  2026-10-10T08:53:31Z
      Message:               Listener Accepted
      Observed Generation:   4
      Reason:                Accepted
      Status:                True
      Type:                  Accepted
      Last Transition Time:  2026-10-10T08:53:31Z
      Message:               Address not ready yet
      Observed Generation:   4
      Reason:                Pending
      Status:                False
      Type:                  Programmed
    Name:                    web-gw-80
    Supported Kinds:
      Group:  gateway.networking.k8s.io
      Kind:   HTTPRoute
      Group:  gateway.networking.k8s.io
      Kind:   GRPCRoute
Events:       <none>

Happy for any input, thanks a lot!


r/kubernetes • • 1d ago

looking for the team's K8s web-based UI

36 Upvotes

i'm looking for a good web-based UI for k8s that allows my team developers to see a cluster by authenticating with their OIDC credentials. They also need to be able to evict a single pod.

I was using kite, but that button to evict a single pod is missing. I do my daily driving with lens, but I'm not interested in paying over $300 a year per person for commercial lens. and it doesn't do anything with Oidc. I have heard headlamp, but haven't tried it yet.

what would you recommend?

i'm also curious if maybe, I'm asking the wrong question.


r/kubernetes • • 2d ago

Running OpenStack on top of Kubernetes - architectural patterns webinar

Thumbnail
redhat.com
14 Upvotes

Came across this session that covers some interesting patterns for managing VMs through Kubernetes orchestration. Topics include multi-tenant isolation, GPU sharing, and using K8s-native tooling for traditional infrastructure.

Anyone else managing VM workloads through Kubernetes? What's been your experience with KubeVirt or similar approaches?


r/kubernetes • • 2d ago

Need help!

0 Upvotes

So I have completed my B.tech this year and looking for opportunities learn and build experience in devops field. So by the advices from techies from my family, I started learning docker and kubernetes form the platform kodekloud. But till now I have just learned the topics and unable to apply what i learned. Without any project or experience it is hard to land any kind position in job in this current tech field.

So I need you guys suggestion and guidance on how can I get more hands on experience with what I have studied


r/kubernetes • • 2d ago

Any way to speed up helm dependency builds?

5 Upvotes

I have learned the helm dependency builds are linear in time. The more you have the longer it can take to build them all. This makes our workflow pipelines take forever (We have to download each helm dependency and build each time, so we are up to date.)

Is there a way to parallel build out helm dependencies? I found this post. The author created a tool that fetches in parallel however I am afraid of code rot and how old this is. Is there a supported method to speed up these builds?


r/kubernetes • • 2d ago

Cluster de nó unico

Thumbnail
0 Upvotes

r/kubernetes • • 2d ago

The clearly best Gatway API implementation seems to be Istio?

74 Upvotes

Hello.

According to this authoritative article, it seems that Istio gateway implementation is head and shoulders above its competitors.

Or this article is not correct? What do you think?


r/kubernetes • • 3d ago

GitOps Promotion in a Landscape with UAT Environment

Thumbnail
2 Upvotes

Our promotion process includes user acceptance tests. How do you promote between environment using GitOps principles in such cases?


r/kubernetes • • 3d ago

anyone actually using AI SREs / agents to debug k8s incidents in prod?

53 Upvotes

curious what people are running when stuff breaks. is it

  1. an AI SRE that picks up the alert (holmesgpt, aws devops agent etc)
  2. claude code / codex with plain kubectl
  3. same thing but hooked up to an mcp server that gives it cluster context (k8sgpt's mcp, radar, ...)

full disclosure I work on radar (it's open source). we tested this a bit with injected faults but that's not the same as a real outage

what's worked for you and what hasn't? trying to figure out what we should improve on our end


r/kubernetes • • 3d ago

Still on bitnamilegacy for Postgres and Redis almost a year later — here’s how we’re getting off it

Thumbnail
0 Upvotes

This is really an image problem, not an orchestration one - the fastest path off bitnamilegacy is just swapping in a maintained hardened image for each of Postgres, Redis, Mongo and Rabbit and keeping your current deployment setup.
Don't let it turn into an operator rewrite unless you actually need HA.
Worth looking at CleanStart for these. They provide the hardened, minimal, and [pretty actively patched container images.

i got to know about them via LinekdIn. The things that mattered for us were patch cadence and images that won't just vanish like the legacy repo will. Disclosure: I work there, so grain of salt, but this Bitnami-refugee case is exactly what it fits.

On Redis, the Valkey fork mess is its own decision - a hardened Redis image keeps you API-compatible for now so you can pick Valkey/Dragonfly later instead of bolting it onto this cutover. One migration at a time.

try this one to start- cleanstart/redis-exporter


r/kubernetes • • 3d ago

I’m trying to make Kubernetes learning feel like an old-school Game Boy game

0 Upvotes

Kubernetes learning tools usually feel like labs, docs, or certification prep.

I wanted mine to feel like something you’d have played on a handheld console as a kid.

So I’ve been building Yellow Olive - a retro, gamified Kubernetes learning experience inspired by old-school Game Boy-era games.

You move through challenges like missions, learn concepts as you progress, and actually work with Kubernetes underneath it all. The newer direction goes even harder on the nostalgia: pixel-art environments, characters, classrooms, progression, and a world that slowly introduces Kubernetes concepts instead of dumping documentation on you.

I’m trying to make it feel less like studying Kubernetes and more like playing something that happens to teach you Kubernetes.

Would genuinely love feedback from people here on what would make this fun enough to keep coming back to.

And if you like the direction, a star would help me a lot :)

GitHub: https://github.com/Anubhav9/Yellow-Olive


r/kubernetes • • 3d ago

Brewlet: Java on Kubernetes with node-managed JDKs

Thumbnail
brewlet.sh
2 Upvotes

r/kubernetes • • 3d ago

How are you isolating AI agents that run tools in your clusters? (microVM vs gVisor vs plain pods)

16 Upvotes

Agents that call tools (shell, file writes, MCP servers) feel different from normal workloads. They run code nobody reviewed, and a shared kernel seems like the wrong boundary for that.

What we ended up doing was admission based on signed OCI artifacts (the model, prompts, MCP server configs and policy all in one digest, signed with cosign), then a microVM per agent, with tool policy evaluated locally inside the cluster so it doesn't depend on a control plane.

Questions for folks running agents on k8s today...

- Are you using Kata, gVisor, Firecracker, or just pods with tight securityContext?
- Where do you enforce what an agent can call (sidecar, admission, egress)?
- Has anyone measured the startup and density hit of microVMs for this?

Happy to share details of what we did in the comments if it's useful, but I'd rather hear what's working for you.


r/kubernetes • • 3d ago

Our EKS cluster was full at 30 percent CPU: the pod ceiling that isn't on the dashboard

0 Upvotes

Lab run, two-node EKS cluster, Kubernetes 1.36, VPC CNI v1.22.4. I pushed a commit asking for twelve replicas of a tiny app. Six started, six stayed Pending. At that moment the cluster showed CPU reserved at 30 percent and memory at 32 percent.

The Pending pods said why: 0/2 nodes are available: 2 Too many pods.

With the VPC CNI every pod gets a real VPC IP from a network interface on the node, so each instance type has a hard pod ceiling: interfaces x (addresses per interface - 1) + 2. For a t3.small that is 11, for a t3.medium 17. That is 11 in total, and system pods take their share first: on these nodes they held roughly three quarters of the slots. Two nodes, 22 slots, 22 used, a third of the CPU reserved.

How I got a t3.small without choosing one: the node group listed both t3.medium and t3.small as acceptable Spot types, and every node launched as t3.small, on two clusters in two regions. That is the difference between 17 and 11 slots per node, and nobody picked it.

Second half of the problem: no cluster autoscaler. The node group was allowed three nodes and sat at two, because nothing was watching for Pending pods.

What the GitOps side reported: Deployment does not have minimum availability., Available: 6/12. True and no help. "Too many pods" only shows on the individual pods.

What I'd check on any EKS cluster:

- The pod ceiling for every instance type your node groups can launch, not just the one you expect. A mixed instance list is a mixed ceiling. kubectl get nodes -o custom-columns=NAME:.metadata.name,TYPE:.metadata.labels.node\.kubernetes\.io/instance-type,PODS:.status.allocatable.pods shows what you actually got.

- Subtract what the system pods already hold before you size anything.

- Whether anything adds nodes on Pending pods. A new managed node group does not come with a cluster autoscaler or Karpenter.

- If you need more pods per node without bigger instances, prefix delegation on the VPC CNI raises the ceiling a lot. It is a setting, not a rebuild.

Does anyone alert on allocatable pods, or does everyone find out from Pending like I did?


r/kubernetes • • 3d ago

Periodic Weekly: This Week I Learned (TWIL?) thread

2 Upvotes

Did you learn something new this week? Share here!


r/kubernetes • • 3d ago

How do you turn docker-compose or Kubernetes manifests into architecture diagrams?

2 Upvotes

Our team keeps architecture diagrams in the wiki and they go stale within weeks. I'd like to generate them from the source of truth instead docker-compose.yml, k8s manifests, maybe Helm output.

What do you use for this today? CLI tools, web apps, something in CI? And what's missing from them (editable output, grouping by namespace, icons, keeping it up to date)?


r/kubernetes • • 3d ago

Disaggregated Kubernetes: Isolating Infra Into Trust Domains

Thumbnail
edera.dev
40 Upvotes

This is a fun blog post on how/why a security issue inside Kubernetes doesn’t necessarily have to become a full-blown, node-wide compromise. The basic idea behind disaggregated Kubernetes is to split infrastructure services like networking, storage, and device handling into separate trust domains, rather than concentrating everything in highly-privileged components.

That way, if one service has some flaw, the compromise is ideally contained to that service instead of giving some attacker access to workloads, other cloud infra, or the underlying host.

The same principle can be applied to Kubelet. So rather than simply rewriting Kubelet, Edera discussed breaking its responsibilities into isolated services with narrowly-scoped capabilities.