03 Aug 2026

feedKubernetes Blog

Gateway API v1.6: TCPRoute and UDPRoute Graduate to Standard

Gateway API logo

The Kubernetes SIG Network community is thrilled to share the release of Gateway API v1.6.0, which was released on June 30th of this year!

Gateway API has become the standard for modern, role-oriented, and expressive service networking in Kubernetes. In previous releases, Gateway API established a production-grade foundation for HTTP and TLS layer 7 traffic. With version 1.6.0, Gateway API takes a major step forward by expanding standard layer 4 protocol routing and introducing cleaner API boundaries for experimental innovation.

Here is a quick summary of what's new in Gateway API v1.6.0:

Let's dive into the details!

TCPRoute and UDPRoute graduate to Standard

Leads: Nick Young, Ricardo Katz and Zac Nixon

Until now, Gateway API only offered a stable routing model for HTTP and TLS traffic. Workloads that speak a raw protocol over TCP or UDP - databases, DNS, VoIP, gaming, IoT telemetry - had no portable way to plug into a Gateway. Users either fell back to a plain Kubernetes Service, or to an implementation-specific CRD that doesn't travel between Gateway controllers.

TCPRoute and UDPRoute close that gap: they route traffic to backends based on protocol and port alone, no L7 awareness required. With this release, both have graduated from the Experimental channel to Standard, and moved to the v1 API version. The v1alpha2 version of each was deprecated as of the v1.6 release, and will be removed in a future release.

How it works

A Gateway needs a listener that allows TCPRoute attachment:

apiVersion: gateway.networking.k8s.io/v1
kind: Gateway
metadata:
 name: example-gateway
spec:
 gatewayClassName: example-gateway-class
 listeners:
 - name: foo
 protocol: TCP
 port: 12345
 allowedRoutes:
 kinds:
 - kind: TCPRoute

A TCPRoute then attaches to that listener and forwards traffic to a backend:

apiVersion: gateway.networking.k8s.io/v1
kind: TCPRoute
metadata:
 name: tcp-app
spec:
 parentRefs:
 - name: example-gateway
 sectionName: foo
 rules:
 - backendRefs:
 - name: my-foo-service
 port: 6000

Traffic arriving on the Gateway's port 12345 is proxied to the endpoints of my-foo-service on port 6000. Omitting sectionName and port from parentRefs attaches the route to every TCP listener on the Gateway instead of a single one.

UDPRoute follows the same pattern; swap the listener protocol and the route kind:

apiVersion: gateway.networking.k8s.io/v1
kind: Gateway
metadata:
 name: example-gateway
spec:
 gatewayClassName: example-gateway-class
 listeners:
 - name: foo
 protocol: UDP
 port: 12345
 allowedRoutes:
 kinds:
 - kind: UDPRoute
---
apiVersion: gateway.networking.k8s.io/v1
kind: UDPRoute
metadata:
 name: udp-app
spec:
 parentRefs:
 - name: example-gateway
 sectionName: foo
 rules:
 - backendRefs:
 - name: my-foo-service
 port: 6000

XBackend arrives in Experimental

Leads: Keith Mattix II

Gateway API v1.6 introduces the new XBackend resource, which is a general-purpose decorator for Service (and other backend types) within Gateway API.

The Service resource is an amazing, stable, and flexible object, but that comes with some costs: The flexibility creates a lot of edge cases that Gateway API needs to handle, and the stability makes it impossible to add new concepts to Service.

The XBackend resource builds on the ideas in the upstream EndpointSelector KEP, to add a Gateway API-native object that still targets the backend app, while allowing the community to extend it to handle use cases that are difficult or dangerous to handle with Service.

The first version of XBackend includes support for ExternalHostname destinations, which are ruled out from Service support in Gateway API because of the possibility of confused deputy attacks.

For XBackend, this support is an Extended/Optional feature, allowing implementations and users to opt in once they understand the security tradeoffs.

This support is very useful for egress use cases (which are most commonly used for cluster-hosted agentic workloads), which the community is also working towards formalizing in GEPs about Gateways for Egress (work in progress, stay tuned!)

The XBackend API is experimental and its behavior can change, do not assume it is ready for production

An example of a Gateway with an ExternalName backend that can be used for egress to a cloud AI API is as follows:


# Gateway-level TLS remains authoritative for incoming connections
apiVersion: gateway.networking.k8s.io/v1
kind: Gateway
spec:
 listeners:
 - name: https
 protocol: HTTPS
 tls:
 certificateRefs:
 - name: gateway-cert
---
# Backend resource for external destination
apiVersion: gateway.networking.x-k8s.io/v1alpha1
kind: XBackend
metadata:
 name: ai-provider-api
 namespace: ai-apps
spec:
 type: ExternalHostname
 externalHostname:
 hostname: api.ai-provider.com

---
# HTTPRoute referencing XBackend
apiVersion: gateway.networking.k8s.io/v1
kind: HTTPRoute
spec:
 rules:
 - backendRefs:
 - name: ai-provider-api
 kind: XBackend
 group: gateway.networking.x-k8s.io

The community is also working on moving Session Persistence config from XBackendTrafficPolicy into XBackend, along with other use cases like retries, TLS origination and similar config that is useful to be able to configure per-application rather than per-Route.

Experimental resources move off the standard API group

Previously, experimental resources shared the same API group as standard ones - gateway.networking.k8s.io - distinguished only by a v1alpha2-style version. TCPRoute and UDPRoute were the last resources to graduate under that scheme.

Going forward, new experimental resources are defined in a separate group, gateway.networking.x-k8s.io, and the names of their API types get an X prefix - for example XBackend and XMesh. When one of these graduates to Standard, it's renamed into the gateway.networking.k8s.io group and drops the X prefix, the same way XMesh is expected to become Mesh.

This separation makes the experimental/standard boundary explicit at the API group level, rather than relying on version strings alone.

What's next & getting involved

The graduation of TCPRoute and UDPRoute to Standard marks an essential milestone in making Gateway API a complete, universal ingress and mesh networking API for Kubernetes workloads across layer 4 and layer 7 protocols.

Try it out

You can start using Gateway API v1.6.0 today with your favorite Gateway controller implementation:

Gateway API relies on an extensive conformance test suite to ensure consistent, portable behavior across all implementations. Here is a list of the implementations that are conforment with v1.6 on the day we published the article:

Get involved

Gateway API is an open, community-driven project built under Kubernetes SIG Network. We welcome contributions, feedback, and participation from everyone!

Acknowledgments

A huge thank you to all the contributors, reviewers, maintainers, and implementation authors whose hard work made Gateway API v1.6.0 possible!

03 Aug 2026 4:00pm GMT

31 Jul 2026

feedKubernetes Blog

Kubernetes v1.37 Sneak Peek

As we get closer to the release date for Kubernetes v1.37, the project develops and matures, features may be deprecated, removed, or replaced with better ones for the project's overall health. This blog outlines some of the planned changes for the Kubernetes v1.37 release that the release team feels you should be aware of for the continued maintenance of your Kubernetes environment and keeping up to date with the latest changes. The information below reflects the current status of the v1.37 release and may change before the actual release date.

Deprecations and removals for Kubernetes v1.37

Kubectl: kubectl run --filename/-f to be deprecated

The --filename (or -f) flag for kubectl run is being deprecated as the generated pod is always built purely from CLI arguments like NAME and --image.

See kubernetes/kubernetes#138671 for the original issue and discussion.

Kubelet: Static Pods can no longer reference Secrets or ConfigMaps

Static Pods were never meant to read API resources directly, since they aren't created through the API server - but a bug let them reference Secrets or ConfigMaps via fields like configMapRef or secretRef. That bug is now fixed: as of v1.37 these references are strictly prohibited, and the PreventStaticPodAPIReferences feature gate that previously let you opt out of the restriction has been removed.

See kubernetes/kubernetes#140226 for the original issue and discussion.

Deprecating kube-proxy's support for ipvs mode

kube-proxy support for ipvs mode was introduced in v1.8 to resolve iptables performance bottlenecks. However, since the kernel ipvs API alone cannot fully implement Kubernetes Services, ipvs mode continues to use iptables underneath (KEP-3866, "The ipvs mode of kube-proxy will not save us").

Clusters running kube-proxy in ipvs mode (or mode: ipvs in KubeProxyConfiguration) would now be logging a deprecation warning on startup. The deprecation timeline looks like this:

kubectl -n kube-system get configmap kube-proxy -o jsonpath='{.data.config\.conf}' | grep 'mode:'

To understand the rationale behind this deprecation, see KEP-5495: Deprecate ipvs mode in kube-proxy.

Ongoing major changes

Future removal of cgroup v1 support

As modern Linux distributions and container runtimes use cgroup v2 as the default, support for the legacy cgroup v1 is officially being phased out. Since the v1.35 release, the failCgroupV1 setting has defaulted to true. Consequently, the kubelet will fail to initialize on any nodes that still rely on cgroup v1 unless an explicit configuration override is applied.

apiVersion: kubelet.config.k8s.io/v1beta1
kind: KubeletConfiguration
failCgroupV1: false # temporary override

Using this override should be considered a short-term fix. Advanced resource management capabilities, such as In-Place Pod Resizing and Tiered Memory Protection, depend entirely on cgroup v2. While the override remains available in Kubernetes v1.37, users are encouraged to migrate to cgroup v2, as support for cgroup v1 is planned to be removed in a future release.

To learn more about this deprecation, refer to KEP-5573: Remove cgroup v1 support.

Breaking changes in Kubernetes v1.37

SELinux volume relabeling ("SELinuxMount") graduates to GA

SELinuxMount is expected to reach GA and be enabled by default in v1.37. Volumes would then be mounted with -o context=<label> (the mount option default) instead of being recursively relabeled, but only when the volume's CSI driver opts in via a CSIDriver that sets .spec seLinuxMount: true.

Because a single mount can only hold one SELinux context, pods with different SELinux labels sharing a volume on the same node (which previously coexisted under recursive relabeling) may now fail to start. To retain the previous recursive behavior for a specific workload, set seLinuxChangePolicy: Recursive in the Pod spec.

Clusters without SELinux enabled see no effect at all. To learn more, check SELinux Volume Label Changes goes GA (and likely implications in v1.37)

Featured enhancements of Kubernetes v1.37

Metrics API goes GA

The metrics.k8s.io API is expected to graduate to Stable (GA) in Kubernetes v1.37 after spending nearly nine years in Beta. The API provides a standard way to retrieve CPU and memory usage for pods and nodes, powering widely used Kubernetes features such as the Horizontal Pod Autoscaler (HPA) and commands like kubectl top.

This graduation recognizes the API's stability and widespread adoption, with no functional changes expected. Both v1 and v1beta1 will remain usable during the transition, enabling developers to adopt the stable API at their own pace without breaking existing workflows.

To learn more about this enhancement, refer to KEP-5207: metrics.k8s.io API definition.

Kubelet in UserNS a.k.a. Rootless Mode

Traditionally, Kubernetes node components such as the kubelet run with root privileges on the host. While necessary for many deployments, this also means that a vulnerability in one of these components could potentially have a greater impact on the underlying system.

With Kubernetes v1.37, kubelet in User Namespace (Rootless Mode) is expected to graduate to Beta. This enhancement allows Kubernetes node components to run inside a Linux user namespace as an unprivileged user on the host while still behaving as root within the namespace. By reducing the need for host-level root privileges, it adds an extra layer of isolation and helps limit the impact of potential vulnerabilities affecting node components.

To learn more about this enhancement, refer to KEP-2033: Kubelet in UserNS(aka Rootless Mode).

Volume health monitor

Historically, Kubernetes has lacked an API for CSI drivers to report storage failures, which become evident only through failed mounts or hung I/O. Since remediation controllers had nothing machine-readable to act upon, the only way to figure out the root cause behind this failure was to cross-reference Kubernetes objects alongside external vendor dashboards.

In Kubernetes v1.37, this KEP resets graduation to Alpha after an initial implementation in v1.21 and introduces four new CSI RPCs. The controller plugin reports the health of storage volumes using ControllerListVolumeHealth (lists unhealthy volumes) and ControllerGetVolumeHealth (checks a specific volume). A controller-side health monitor polls these CSI controllers and stores the results in PersistentVolumeClaim.status.healthStatus.

On the node side, the kubelet calls NodeGetVolumeHealth to obtain the health of individual volumes on that node and records it in Pod.status.volumeHealth, while NodeGetStorageHealth reports the health of the drivers registered to a node in CSINode.status.storageHealth.

The error vocabulary is kept simple, extensible, and machine-parsable (Inaccessible, Degraded, etc.), with further driver-specific elaboration available via reason and message. Finally, the controller-side and node-side reports are kept independent and are hence displayed separately, providing a more holistic view of storage health to consumers.

To learn more about this enhancement, refer to KEP-1432: Volume Health Monitor.

Want to know more?

New features and deprecations are also announced in the Kubernetes release notes. We will formally announce what's new in Kubernetes v1.37 as part of the CHANGELOG for that release.

Kubernetes v1.37 release is planned for Wednesday, August 26th, 2026. Stay tuned for updates!

You can see the announcements of changes in the release notes for:

Get involved

The simplest way to get involved with Kubernetes is by joining one of the many Special Interest Groups (SIGs) that align with your interests.

If you don't know where to start, join our monthly New Contributor Orientations where we teach the community how the project is structured, and we'll guide you on how to make your first contribution to the project.

31 Jul 2026 4:00pm GMT

29 Jul 2026

feedKubernetes Blog

How the controller-runtime Cache Actually Works, and Why Your Controller Does Not Crash the API Server

Caution:

Some of the technical detail in this article is not accurate. We are reviewing it and preparing corrections. Until then, check what you read here against the controller-runtime documentation.

Kubernetes has long been the default platform for distributed workloads, and writing your own controller for it is now a matter of a few hours. The common path - Golang, using kubebuilder on top of controller-runtime - gives you a project scaffold, types, and a reconciler. For typical scenarios that is more than enough. But as soon as load grows or the controller starts behaving in ways you did not expect, a whole class of edge cases shows up. Most of them trace back to the same root cause: a fuzzy mental model of how controller-runtime works inside. If you write Kubernetes controllers in Go, this article should help you build a coherent picture and avoid expensive surprises in production.

This article walks through the internals of controller-runtime and, along the way, shows which architectural decisions are baked into Kubernetes itself. The starting point is how controllers actually read objects from the Kubernetes API.

A common misconception goes like this: r.Get() inside Reconcile queries kube-apiserver directly; r.List() returns a fresh, live view of the world; and after r.Update() you can re-read the object and immediately see the new state. In practice the model is the opposite: controller-runtime operates against a local copy of the data populated through list + watch. Reads inside a reconciler cost almost nothing and do not load the control plane even at hundreds of calls per second - but the price of this design is that a controller can quietly consume gigabytes of memory, perform hidden O(n) scans, and regularly trip over stale reads.

This post is aimed at engineers who already write controllers in Go with controller-runtime but want to consolidate the pieces into a single mental model rather than carry around a bag of isolated observations. The focus is the practical impact on production clusters: memory, network traffic, read consistency, and reconciler behavior.

TL;DR

If you take only one idea from this article, take this:

r.Get() and r.List() inside a reconciler typically do not read from the API server. They read from a local in-memory cache, which the manager warms up with list and then keeps current through watch.

Almost every other property of the system follows from that one fact:

The rest of the article unpacks why this is so and how the model is wired underneath.

A bit of context: what a reconciliation loop is

To avoid arguments about terminology, start with the basic model.

A controller in Kubernetes lives inside a reconciliation loop: it continuously compares the desired state of an object with the actual state and tries to bring one in line with the other. The idea is described in the original architectural notes on Kubernetes. In practice it looks like this:

What matters here is not that the controller "does something" - it is where it learns about changes from and where it reads state from. That is exactly where the cache comes in.

On a live cluster, the easiest way to see this in action is:

kubectl get pods --watch

In watch mode, kubectl subscribes to the same event stream that controllers consume. You create or delete a Pod and you see not a single "final" object but a chain of states: the scheduler assigns a node, the kubelet updates status, other controllers contribute their changes. Kubernetes controllers do not poll continuously - they consume an event stream and maintain a local state that is kept current.

For a visual walkthrough, see Reconciliation loop pattern in visual representation, a talk that shows how the reconciliation loop plays out on a real Pod and the states it passes through.

Why the cache exists in controller-runtime at all

Imagine the simplest possible controller:

func (r *Reconciler) Reconcile(ctx context.Context, req ctrl.Request) (ctrl.Result, error) {
 var pod corev1.Pod
 if err := r.Get(ctx, req.NamespacedName, &pod); err != nil {
 return ctrl.Result{}, err
 }
 // ... meaningful logic ...
}

Looks straightforward. But what happens when you call r.Get? Does it fire an HTTP request at the API server? If it did, picture the scene: a dozen controllers, each issuing a get and a list per reconcile, with hundreds of reconciles per second. The API server and etcd would be writing each other farewell letters within minutes.

To prevent that, Kubernetes was built around a watch model rather than polling from the very beginning. The standard mechanism works like this: a client issues list once, gets a snapshot of the slice of the world it cares about, then subscribes to a stream of changes via watch and keeps a local copy current. Everything happens over a single long-lived HTTP connection, with no "what is in the world right now?" loop.

This idea has lived in client-go since the very first controllers in kube-controller-manager. controller-runtime wraps it in a friendly framework so that you do not have to glue Reflector, DeltaFIFO, and Indexer together yourself (more on those below).

So when people talk about "the controller-runtime cache", they are not talking about a clever optimization. They are describing the foundation of the entire model: you read from memory, you write to the API server, and you receive feedback through a watch.

The rest of this article walks through how each piece is wired up.

Glossary

A few terms collected up front, so you do not have to jump back and forth later. Skim or skip if any of them are already familiar.

With those in hand, you can dive in.

Anatomy: what lives under the cache package

If you peek into sigs.k8s.io/controller-runtime/pkg/cache, you will see that it is a thin wrapper over k8s.io/client-go/tools/cache. The same primitives that power the rest of Kubernetes live underneath:

At a glance the pipeline looks like this:

A vertical flow chart with seven labeled boxes connected by arrows, from API server at the top to Reconcile at the bottom.

Pipeline diagram: API server to Reflector to DeltaFIFO to Indexer to Event handlers

Now walk through each link.

Reflector and resourceVersion

The Reflector is the only component that talks to the API server directly. It has exactly two jobs: do a single list at startup, then keep a watch open from there on.

This is where the resourceVersion earns its keep. Along with the list of objects, the API server returns the version at which the snapshot was produced. The Reflector then says to the API server, "open a watch from version X", and receives a stream of events for everything that happened after that version. That is the basis of consistency: there is no risk of missing an event between list and watch, because watch resumes exactly at the point where list ended.

If the connection drops, the Reflector reconnects with the last known resourceVersion. If the API server replies with 410 Gone ("that version is no longer in the history, you are too far behind"), the Reflector performs a fresh list and starts over. This is called a relist, and it does not happen on a schedule - only in those failure scenarios.

DeltaFIFO: a queue of deltas

This piece is worth pausing on. DeltaFIFO is the buffer between the Reflector and the rest of the informer. Its input is a stream of events from the API server; its output is the same events, but grouped by key and in strict order.

More precisely, DeltaFIFO solves three problems:

  1. It preserves order. Whatever stream of changes flows in for default/my-deploy, the consumer sees the same ordering the API server delivered.
  2. It groups by key. All deltas for a single namespace/name accumulate in one slot. Pop() returns not a single delta but a slice of every delta accumulated under that key - the consumer sees, in one shot, everything that has happened to the object since the last call.
  3. It deduplicates selectively. The built-in dedupDeltas function collapses consecutive Deleted deltas for the same key, so two delete events do not turn into two separate processing rounds.

An important caveat: DeltaFIFO does not merge consecutive Added or consecutive Updated deltas. Collapsing every intermediate state into a single final one is, in general, not its job.

A worked example. Suppose three events for object default/my-deploy arrive in quick succession:

  1. Added - the Deployment is created (say, with spec.replicas=1).
  2. Updated - somebody bumps spec.replicas to 2.
  3. Updated - and immediately to 3.

DeltaFIFO places all three deltas into the slot keyed by default/my-deploy. Pop() returns them as a single slice, and sharedIndexInformer.HandleDeltas walks through them in order: first OnAdd, then two OnUpdate calls (one for the intermediate 1→2 transition and one for the final 2→3). The event handler runs three times, no shortcuts.

There is per-object deduplication, but not in DeltaFIFO - it lives one layer up, in the controller's workqueue. The mechanic is straightforward: for each delta from DeltaFIFO, the controller's event handler extracts the namespace/name key from the object and enqueues it. Re-inserting the same key silently coalesces with the existing entry; the workqueue does not care about the object itself.

A concrete picture: you create a Pod. Within a second or two a flurry of Updated deltas arrives - the scheduler assigns a node, the kubelet sets Pending, then ContainerCreating, Running, Ready. Five deltas in a row, and the event handler fires on every one of them - but throughout this window the workqueue holds a single entry with the key default/my-pod. By the time Reconcile pops it, the cache already holds the final state, and Reconcile runs once.

So you get two layers with cleanly separated responsibilities:

If you keep that two-layer picture in your head, it becomes clear why a flood of events against a single object barely affects controller throughput - the workqueue absorbs them.

Indexer: the local copy of the cluster

The Indexer (also known as ThreadSafeStore) is the local copy of the cluster. Underneath it is a plain map[string]interface{} keyed by namespace/name, plus a mutex, plus a dictionary of registered indexes (covered in their own section below).

Yes - at heart it is a map in memory. No B-trees, no LSMs. That is precisely why a cache-hit r.Get costs microseconds: it is a map lookup followed by a copy of a Go struct.

SharedIndexInformer and subscriptions

A SharedIndexInformer fuses Reflector, DeltaFIFO, and Indexer together and exposes two interfaces to the rest of the world:

"Outside" here means your controllers. When a controller registers Watches(...), under the hood it asks the informer: "add a handler that, on every change, enqueues the key into my workqueue". The controller's workers then pop keys one at a time and call your Reconcile(ctx, ctrl.Request{NamespacedName: ...}).

The keyword in the name is Shared. The manager creates one informer per GVK, and every controller, webhook, and event source within that manager subscribes to it:

A single Pod informer at the top with three arrows fanning out to two controllers and a webhook, all inside a ctrl.Manager box.

Shared informer diagram: a single list / watch per GVK, feeding multiple subscribers

In other words: an informer is the thing that subscribed to Pods once, holds them locally, and serves every interested party in the process. From the API server's perspective, that is one list and one watch per GVK, regardless of how many reconcilers live inside your process.

What happens at startup and on the very first r.Get

Step by step, here is what happens between the moment the manager starts and the first r.Get inside your reconciler:

  1. The manager's mgr.Start(ctx) brings up every registered informer.
  2. For each GVK, the Reflector performs a full list of every object that falls within your scope.
  3. The list response is loaded into the informer's store, registered indexes are rebuilt, and the informer's HasSynced() flag flips to true.
  4. After that, a watch is opened starting from the resourceVersion returned by list.
  5. Only then does the controller start invoking Reconcile - specifically, once cache.WaitForCacheSync has returned true for every source it owns. Until that point, workers do not drain the workqueue, even if events have already started piling up.

So in controller-runtime, "the reconciler is running but the cache is still empty" is not a state you can ever observe by construction. The warm-up always happens up front, never lazily.

What happens during the first r.Get? Suppose your reconciler contains:

var obj appsv1.Deployment
err := r.Get(ctx, req.NamespacedName, &obj)

Under the hood it boils down to roughly this:

item, exists, err := indexer.GetByKey("default/my-deploy")
if !exists {
 return apierrors.NewNotFound(...)
}
// DeepCopy into obj

No HTTP, no TLS, no protobuf serialization, no etcd. A map lookup, a struct copy, return. Microseconds.

To repeat, because it matters: even the very first Get in the controller's lifetime reads from a fully warmed-up, fully indexed snapshot. There is no "first time slow, then fast".

Note: This applies specifically to mgr.GetClient(). If for some reason you need to read objects before mgr.Start() (for example, during initialization), use mgr.GetAPIReader(), which goes straight to the API server. More on this later.

Client ≠ Cache: read from memory, write to the API server

Another point that often gets lost. client.Client in controller-runtime is a composite object:

This is not a hack - it is a deliberate design choice:

It is worth dwelling on "should be exact". This is where resourceVersion shows up again.

When you read an object from the cache, you do not get its current state in etcd - you get the state as the Reflector last observed it. That state carries a resourceVersion. You then mutate the object and call r.Update(ctx, &obj). The request goes to the API server right now, and the API server checks:

This is optimistic concurrency control. No real locks are taken; everybody writes in parallel; but only one of the racing Update calls wins - the one that arrives with the current version. Everyone else gets a 409 and is expected to re-read and try again.

Why does this matter for the cache? If you naively send a PUT with "your" resourceVersion from the cache and somebody has updated the object since you read it, you will get 409. That is not a bug. It is exactly the protection the system is supposed to give you. Writing without the resourceVersion check (via Patch without an optimistic lock, or via Server-Side Apply) is also possible, but that is a separate conversation.

The "write → visibility" cycle now looks like this:

A vertical diagram showing how a write travels from user code through the API server and back into the controller's cache via a watch event.

Write visibility diagram: client.Update to API server to watch event to cache

Between "you executed Update" and "the cache reflects the new state" there is a microscopic window, on the order of milliseconds. Inside that window, an r.Get for the same object returns the previous version. The next section is essentially a list of mistakes that grow out of that window.

Common mistakes that everyone makes

Mistake 1: expecting read-after-write

A familiar pattern:

obj.Spec.Replicas = ptr.To(int32(5))
if err := r.Update(ctx, &obj); err != nil {
 return ctrl.Result{}, err
}

// re-read and confirm it is now 5
var fresh appsv1.Deployment
_ = r.Get(ctx, key, &fresh)
fmt.Println(*fresh.Spec.Replicas) // surprise: 3

This is not a controller-runtime bug. It is a property of an eventually consistent system: the cache catches up asynchronously, through the watch.

The right pattern is to never rely on instant freshness. Reconcile must be idempotent and must always look at the current state. If it does not match the desired state, the next reconcile fixes it. You do not need to "wait 100ms" or "re-trigger". You need to write the logic so that one or two extra invocations break nothing.

If you genuinely need guaranteed freshness - for example, in a validating webhook where you cannot afford to act on stale state - that is what APIReader is for. More on this shortly.

Mistake 2: DeepCopy and who owns the memory

To make sense of this, a quick word on event mechanics inside a controller. When you register a source via Watches(...), two layers sit between the indexer and your Reconcile:

Here is the critical part. Predicates and handlers receive the same objects that live in the informer's shared store. The same *corev1.Pod is seen by every controller subscribed to Pods.

Because Go has no immutable structs, nothing prevents you from doing pod.Labels["foo"] = "bar" directly inside a handler. Historically, Get and List returned a pointer into the store as well, with predictable consequences: somebody patched a status "for convenience" in one controller and broke the world view of an unrelated controller next door.

Today, controller-runtime performs a DeepCopy on Get and List by default. The simple rule:

A concrete review heuristic: if predicate.Funcs{UpdateFunc: ...} or handler.EnqueueRequestsFromMapFunc(...) contains expressions like e.ObjectNew.SetLabels(...) or obj.Status.X = Y, stop and ask whether a DeepCopy is missing before that mutation.

Mistake 3: resync is not relist

An informer has a resyncPeriod parameter (10 hours by default in controller-runtime), and many people read it as "rebuild the cache from the API server every N hours".

It does not. A resync does not perform a list. It re-emits everything currently in the indexer back through DeltaFIFO as Sync deltas, and the informer processes them as usual, calling OnUpdate(old, old) for each object. This gives a controller that has somehow missed its reconcile window (a stuck worker, a dropped handler) a chance to see the world again. It generates no traffic to the API server.

A real relist happens only in two cases: when the watch died with 410 Gone, and when you explicitly recreate the informer.

Mistake 4: do not confuse RequeueAfter with a timer

A small note that often saves time. Sometimes you want to wait inside a reconciler - "we just called the provider's API; if it is not ready yet, retry in a minute". The temptation is to spin up time.Sleep or your own goroutine.

Resist it. controller-runtime already provides a built-in mechanism:

return ctrl.Result{RequeueAfter: 30 * time.Second}, nil

The controller puts your req back into the workqueue with a delayed trigger 30 seconds out. If a real event for the same object arrives within that window, the reconcile fires immediately, without waiting for the timer (the key is deduplicated in the queue). This is both cheaper and more correct than a hand-rolled timer: you do not hold a worker, and you do not risk missing a real event.

There is also ctrl.Result{Requeue: true} - enqueue immediately, subject to the rate limiter.

cache + index = almost SQL

Now you get to what is, arguably, the most useful capability of the cache - and the one most controllers leave unused.

By default, a List from the cache looks like this:

var pods corev1.PodList
_ = r.List(ctx, &pods)
for _, p := range pods.Items {
 if p.Spec.NodeName == "node-1" {
 // do something
 }
}

It works - until the cluster has 50,000 Pods and reconciles run hundreds of times per second, at which point the controller is shuffling the same half-gigabyte of pointers back and forth on every trigger. O(n) per reconcile.

The Indexer in client-go can do much better. You declare up front which field you want to index on:

// Index by spec.nodeName for Pods
if err := mgr.GetFieldIndexer().IndexField(
 ctx,
 &corev1.Pod{},
 "spec.nodeName",
 func(obj client.Object) []string {
 pod := obj.(*corev1.Pod)
 if pod.Spec.NodeName == "" {
 return nil
 }
 return []string{pod.Spec.NodeName}
 },
); err != nil {
 return err
}

Two things about that call are worth making explicit, because the tidy example hides them behind a convention.

The index name is arbitrary. That second argument, "spec.nodeName", is only a string key the index is registered under. controller-runtime does not parse it as JSONPath and does not check it against the object's schema - you could write "by-node" or "xyzzy" and it would behave identically. The only rule is that the exact same string comes back in MatchingFields at query time. Naming the index after the field it happens to read is a readability convention, nothing more.

The indexed value is computed, not read. The function returns whatever strings you build; they need not be the verbatim contents of any single field. You can lowercase a value, join several fields into one composite key, bucket a timestamp (the time-bucket trick below does exactly this), or emit a string that appears nowhere in the object literally. Whatever the function returns becomes a key in the inverted dictionary, and a MatchingFields lookup for that exact key is what finds the objects again. The only constraint is that the value has to be derivable from the object you are indexing.

What is an inverted index? The term comes from search engines. Normally you have documents and each document has a list of words in it. "Inverted" means the relationship is flipped: a dictionary in which the key is a word and the value is the list of documents that contain it. Same idea here: the key is the value of a field (for example, node-1), and the value is the list of object keys whose field has that value:

map["node-1"] = {"default/pod-a", "kube-system/pod-b", ...}
map["node-2"] = {"default/pod-c", ...}

What the indexer does:

And now you can write:

var pods corev1.PodList
_ = r.List(ctx, &pods,
 client.MatchingFields{"spec.nodeName": "node-1"},
)

This is not "fetch the full list, then filter". It is a lookup in the inverted index → a ready set of keys → a fetch of the corresponding objects. A different code path entirely.

The comparison to SQL is more accurate than it might look at first:

SQL controller-runtime
CREATE INDEX idx_node ON pods(node_name) IndexField(&Pod{}, "spec.nodeName", fn)
SELECT * FROM pods WHERE node_name = 'node-1' List(&pods, MatchingFields{"spec.nodeName": "node-1"})
SELECT * FROM obj WHERE owner_uid = $1 List(&list, MatchingFields{"metadata.ownerReferences.uid": uid}) (requires an IndexField for that field)

Note the last row: MatchingFields does not make magic out of thin air. For every field you want to look up via MatchingFields you need a corresponding IndexField registered during manager setup. Without one, controller-runtime rejects the query and returns an error.

A few things worth keeping in mind:

Note: An index is built at registration time and is populated as part of the initial list. By the time the first Reconcile runs, both Get and List with MatchingFields work correctly - the index is not built lazily.

Selective cache: do not pull the whole cluster into your controller

By default, an informer pulls every object of its type from every namespace. For Pod, Secret, ConfigMap, and Event in a large cluster, that is a multi-gigabyte surprise delivered on the first list at startup.

It hurts especially with:

In controller-runtime, caching policy lives in cache.Options, passed when constructing the manager:

mgr, err := ctrl.NewManager(cfg, ctrl.Options{
 Cache: cache.Options{
 ByObject: map[client.Object]cache.ByObject{
 // Cache Secrets only from your own namespace, and only by label
 &corev1.Secret{}: {
 Namespaces: map[string]cache.Config{
 "my-controller": {},
 },
 Label: labels.SelectorFromSet(labels.Set{
 "app.kubernetes.io/managed-by": "my-controller",
 }),
 },
 // Cache all Pods, but trim noise on the way into the store
 &corev1.Pod{}: {
 Transform: func(obj any) (any, error) {
 pod := obj.(*corev1.Pod)
 pod.ManagedFields = nil
 return pod, nil
 },
 },
 },
 },
})

A subtle point: this is a manager-level setting and it affects every controller in the process that reads the corresponding type. If you narrow the cache for Secrets to a single namespace and another controller in the same binary needs all secrets in the cluster, that controller will not see them. Before you tighten the scope, audit who else is reading the type.

A short tour of the options:

Caveat: A selector limits what is cached, not what exists. If an object does not match your selector, then as far as your controller is concerned, it does not exist in either Get or List. This bites people: somebody mislabels a single Secret and then spends half a day figuring out why their controller "cannot see it".

Metadata-only: when spec and data are not needed

A separate pattern: you need to know that an object exists, but you do not need its spec or data. Typical examples: a controller that waits for a Secret with a particular name to appear but never reads it; one that counts PersistentVolume objects by the topology.kubernetes.io/zone label; one that reacts to ConfigMap objects in a namespace by name and does not care about contents.

Caveat: PartialObjectMetadata by definition gives you nothing from spec or status - only ObjectMeta. So you cannot filter through it on spec fields (such as a PersistentVolume's storageClassName or a Pod's nodeName); those fields do not exist in the local copy. Everything covered by metadata-only is labels, annotations, ownerReferences, finalizers, creationTimestamp, and the rest of metadata.

For this case there is PartialObjectMetadata:

var list metav1.PartialObjectMetadataList
// Note: Kind is the singular ("Secret"), not "SecretList".
// controller-runtime infers the list shape from the variable type.
list.SetGroupVersionKind(schema.GroupVersionKind{
 Group: "",
 Version: "v1",
 Kind: "Secret",
})
if err := r.List(ctx, &list, client.InNamespace("my-ns")); err != nil {
 return err
}

Under the hood this is a separate watch that asks the API server for metadata only. The store keeps such objects without Data, Spec, or Status - only ObjectMeta. For Secrets the memory difference can reach an order of magnitude.

APIReader: when the cache is not enough

mgr.GetAPIReader() returns a client.Reader that goes straight to the API server, around the cache. When you actually need it:

The price is a real network request. One thing to avoid: do not build "look in the cache, and if missing, fall back to the API" logic. That is exactly the split-brain pattern the cache is meant to protect you from.

Disabling the cache for a type entirely

If you do not need a local cache for a given type at all - say, the type is "fat", read rarely, and the list + watch overhead is not worth paying - you can tell the manager not to cache it. This is configured through client.Options.Cache.DisableFor:

mgr, err := ctrl.NewManager(cfg, ctrl.Options{
 Client: client.Options{
 Cache: &client.CacheOptions{
 DisableFor: []client.Object{
 &corev1.Secret{},
 },
 },
 },
})

With this configuration, mgr.GetClient().Get(...) and List(...) for Secret go straight to the API server, bypassing the cache. No informer is started for that type, which means no list at startup and no permanent memory pressure from a store. This is a more radical alternative to APIReader: where APIReader is reached for ad hoc, individual requests, DisableFor turns the cache off for the type wholesale.

Real-world projects use this. Several established CNCF operators disable caching on Secrets, both to save memory and to avoid hammering the API server with a large list at startup.

Aside: If you want to avoid a watch on the API server entirely, you can feed the controller events from a source of your own design, bypassing list + watch. In controller-runtime this is done with WatchesRawSource / source.Channel: you can wire the controller to events from any place - an internal queue, a kubelet, a custom watch. Niche, but a perfectly valid pattern when the API server should not be touched.

Good practices

A short checklist worth running through before you ship a controller into a live cluster:

Wrapping up

In one breath:

And the single sentence to remember: r.Get inside a reconciler does not call the API server. Ever. Not even the first time. Once that becomes a reflex, half the questions on controller code reviews answer themselves.

29 Jul 2026 6:00pm GMT