06 Oct 2026
Kubernetes Blog
The Shift to cgroup v2 in Kubernetes: What You Need to Know
In Linux, cgroups (control groups) are a kernel feature used for managing system resources. Kubernetes uses cgroups to allocate resources like CPU and memory to containers, ensuring that applications run smoothly without interfering with each other. With the release of Kubernetes v1.31, support for v1 cgroup management moved into maintenance mode. Support for v2 cgroup management has been stable since Kubernetes v1.25.
Compared with cgroup v1, cgroup v2 provides a single unified hierarchy, a more consistent interface, and a stronger foundation for resource isolation and modern resource-management features.
Deprecation of cgroup v1
Kubernetes has deprecated cgroup v1. Starting with Kubernetes v1.35, failCgroupV1 defaults to true, so the kubelet does not start on a cgroup v1 node by default. Administrators can temporarily set failCgroupV1: false in the kubelet configuration file, but removal will follow the Kubernetes deprecation policy. Further removal work is tracked in KEP-5573: Remove cgroup v1 support.
If you are still on a release older than v1.35, migrate every Linux node to cgroup v2 before upgrading, or plan to set the temporary failCgroupV1: false override. If you are already on v1.35 or later, confirm that every Linux node runs cgroup v2 (or that you intentionally keep the override). Under the default configuration, a remaining cgroup v1 node fails during kubelet startup.
For kubeadm-managed clusters, Kubernetes v1.35 also makes this an earlier, stricter check. The SystemVerification preflight check, provided by k8s.io/system-validators, returns an error during kubeadm init, kubeadm join, and kubeadm upgrade when it detects cgroup v1 with kubelet v1.35 or later; with an older kubelet, the check remains a warning. See kubernetes/system-validators#1.12.1 release notes for details.
The top FAQs cover three main areas: why to migrate, the benefits and drawbacks, and key points to keep in mind when using cgroup v2.
Limitations of cgroup v1 and Improvements with cgroup v2
The Linux kernel documentation describes both interfaces:
Let's enumerate some known issues.
active_file memory is not considered available memory
The kubelet treats active_file memory as not reclaimable. For I/O-intensive workloads, a large page cache can therefore make the kubelet report memory pressure and evict Pods. This is a known kubelet issue (kubernetes/kubernetes#43916); migrating to cgroup v2 does not by itself change that calculation. The documented workaround is to set equal memory requests and limits for containers that perform intensive I/O, after measuring an appropriate value.
Memory QoS updates in Kubernetes v1.36
Memory QoS was introduced as an alpha feature in Kubernetes v1.22 and updated in v1.27. It remains alpha in v1.36, but now separates memory throttling from memory reservation and adds tiered memory protection:
Memory QoS is available only on Linux nodes that use cgroup v2. It relies on the cgroup v2 memory controller: memory.high provides throttling, while memory.min and memory.low provide hard and soft protection when tiered reservation is enabled. cgroup v1 cannot provide this protection model.
-
Enabling the
MemoryQoSfeature gate appliesmemory.highthrottling to Burstable containers. The threshold is derived from the request, limit, andmemoryThrottlingFactor(default0.9). -
memoryReservationPolicy: Noneis the default. It does not writememory.minormemory.low. -
memoryReservationPolicy: TieredReservationmaps Guaranteed Pod memory requests tomemory.min(hard protection) and Burstable Pod requests tomemory.low(soft protection). BestEffort Pods receive neither protection. -
The kubelet exposes Alpha metrics for the total
memory.minandmemory.lowreservations on a node. -
Kernel 5.9 or later is recommended. On older kernels,
memory.highreclaim can trigger a known livelock; from v1.36 the kubelet logs a warning when Memory QoS is enabled on an affected kernel.
For example, to opt in to tiered protection:
apiVersion: kubelet.config.k8s.io/v1beta1
kind: KubeletConfiguration
featureGates:
MemoryQoS: true
memoryReservationPolicy: TieredReservation
memoryThrottlingFactor: 0.9
The overall Kubernetes recommendation is not to enable Alpha features in production; however, if you judge that memory QoS with tiered reservations is useful for your platform in production, make sure to test the configuration and account for hard-reserved memory before you enable the TieredReservation feature gate. See the Pod QoS documentation for the current mapping of Kubernetes QoS classes to cgroup v2 controls.
Container-aware OOM handling
On cgroup v2 nodes, the kubelet defaults singleProcessOOMKill to false. It therefore sets memory.oom.group for each container cgroup so that an OOM event kills all processes in that container together, rather than leaving a partially functioning multi-process container. Set singleProcessOOMKill: true only if you need the cgroup v1-compatible behavior where the kernel may kill one process at a time. See the KubeletConfiguration reference for this setting.
This behavior is scoped to a container cgroup, not the whole Pod. Also, cgroup.kill is a separate administrative interface: writing 1 to it sends SIGKILL to every process in that cgroup and its descendants; it does not configure OOM behavior. The cgroup v2 memory controller additionally provides memory.events counters that monitoring systems and userspace OOM managers can observe.
Rootless support
In cgroup v1, delegating controllers to less privileged containers may be dangerous.
Unlike cgroup v1, cgroup v2 officially supports delegation. Most implementations of rootless containers rely on systemd for delegating v2 controllers to non-root users.
This delegation mechanism is separate from Kubernetes Pod user namespaces, which map container users to unprivileged host users. Pod user namespace support graduated to stable in Kubernetes v1.36; check its filesystem, kernel, CRI runtime, and OCI runtime prerequisites before enabling it.
What else?
- eBPF stories:
- In cgroup v1, device access controls are exposed through interface files.
- The cgroup v2 device controller has no interface files, and is implemented on top of cgroup BPF.
- Cilium attaches BPF cgroup programs for socket-based load balancing. Its default cgroup root is
/run/cilium/cgroupv2.
- Pressure Stall Information (PSI) reports CPU, memory, and I/O contention at node, Pod, and container level. On supported clusters, the kubelet exposes PSI by default (
KubeletPSIis stable and locked on). PSI requires cgroup v2, Linux 4.20 or later,CONFIG_PSI=y, and a kernel not booted withpsi=0. The kubelet surfaces the data through the Summary API and/metrics/cadvisor. - When migrating, update software that reads the cgroup filesystem directly. The migration guidance recommends cAdvisor v0.43.0 or later and lists compatible Java, Node.js, and
automaxprocsversions.
CPU weight conversion in newer OCI runtimes
cgroup v1 uses cpu.shares, whereas cgroup v2 uses cpu.weight. Newer OCI runtimes use an improved non-linear conversion that preserves the default priority and gives small CPU requests more usable granularity. The change is implemented in the OCI runtime rather than Kubernetes: it is available in crun v1.23 and runc v1.3.2. After upgrading a runtime, monitoring or policy tools that predict exact cpu.weight values may need updates. Read New Conversion from cgroup v1 CPU Shares to v2 CPU Weight for the formula, examples, and compatibility considerations.
In-place resource updates
In-place Pod vertical scaling graduated to stable in Kubernetes v1.35. Kubernetes v1.36 then [enabled] (/blog/2026/04/30/kubernetes-v1-36-inplace-pod-level-resources-beta/) in-place vertical scaling for Pod-level resources by default, as a Beta feature. The kubelet coordinates changes between the Pod-level and container cgroups so that increases create headroom before container limits grow, while decreases constrain containers before shrinking the Pod-level boundary. Accurate aggregate enforcement for this v1.36 feature requires cgroup v2.
Adopting cgroup version 2
Requirements
Here's what you need to use cgroup v2 with Kubernetes. First up, you need to be using a version of Kubernetes with support for v2 cgroup management; that's been stable since Kubernetes v1.25 and all supported Kubernetes releases include this support.
- You need at least one Linux node; cgroup is a Linux-only concept
- Your OS install must run with cgroup v2 enabled
- The kernel version must be 5.8 or later (5.9 or later is recommended when using memory QoS)
- The container runtime must support cgroup v2. For example:
- containerd v1.4 or later supports cgroup v2; use containerd v2.0 or later for automatic cgroup-driver discovery
- CRI-O v1.20 or later
- The kubelet and the container runtime must both be configured to use the correct cgroup driver. See Configure the kubelet's cgroup driver to match the container runtime cgroup driver.
For now, you can opt back in to use cgroup v1; the Kubernetes project recommends using cgroup v2, but in Kubernetes 1.36 (the current release) the cgroup v1 option remains supported as a fallback. That fallback is scheduled for removal in Kubernetes v1.38. If you are running an older cluster, plan to migrate; if you are setting up a new cluster with Linux nodes, you should prefer cgroup v2. In either case, review both the Kubernetes runtime documentation and (if relevant) the containerd compatibility matrix.
kernel updates around cgroup v2
When Kubernetes was first announced, in 2014, only v1 cgroup existed. Version 2 cgroup management first appeared in Linux kernel 4.5, released in 2016.
- In Linux 4.5, the cgroup v2
io,memory, andpidscontrollers were supported. - Linux 4.15 added support for the cgroup v2
cpucontroller. - Pressure Stall Information (PSI) support began with Linux 4.20.
- The Kubernetes project does not recommend using cgroup v2 with a Linux kernel older than 5.2 due to lack of cgroup-level task freezer support.
- Kubernetes documents 5.8 as the minimum kernel version for cgroup v2; the root cgroup's system-level
cpu.statfile was added in Linux 5.8. - The
memory.highlivelock fix used by Memory QoS is present in Linux 5.9 and later. memory.peakwas added in Linux 5.19.
cgroup driver configuration
Configure the kubelet's cgroup driver to match the container runtime cgroup driver.
If you use kubeadm to manage your cluster, Kubernetes recommends that you use the systemd cgroup driver, because kubeadm manages the kubelet as a systemd service. For other management tooling, check the documentation for the tool you're using to manage your cluster.
If you can pick either option, I recommend using the systemd driver.
Whatever tooling you've chosen, the kubelet automatically tries to detect the runtime's recommended cgroup driver. This automatic detection relies on using a runtime that implements the RuntimeConfig CRI RPC (for example: containerd v2.0+ or CRI-O v1.28+).
If you're using a container runtime that supports cgroup v2 but doesn't support automatic cgroup driver detection, you can manually configure an override by editing the kubelet configuration file. For example:
apiVersion: kubelet.config.k8s.io/v1beta1
kind: KubeletConfiguration
cgroupDriver: systemd
Automatic discovery of the runtime's cgroup driver through the CRI, tracked by KEP-4033, graduated to stable in Kubernetes v1.34. It requires a runtime that implements the RuntimeConfig CRI RPC (containerd v2.0+ or CRI-O v1.28+). When available, the kubelet uses the value reported by the runtime instead of its configured cgroupDriver value.
Tools and commands for troubleshooting
Tools and commands that you should know about cgroups:
stat -fc %T /sys/fs/cgroup/: Check whether cgroup v2 is enabled; it returnscgroup2fs.systemctl list-units 'kube*' --type=sliceor--type=scope: List Kubernetes-related units that systemd currently has in memory.bpftool cgroup list /sys/fs/cgroup/*: List all programs attached to the cgroup CGROUP.systemd-cgls /sys/fs/cgroup/*: Recursively show control group contents.systemd-cgtop: Show top control groups by their resource usage.tree -L 2 -d /sys/fs/cgroup/kubepods.slice: Show Pods' related cgroups directories.
How to check if a Pod CPU or memory limit is successfully applied to the cgroup file?
Work from the API object down to the node. Identify the node, then compare desired resources in the spec with the enacted values in status (in-place resize can leave those out of sync):
kubectl get pod <pod-name> -n <namespace> \
-o jsonpath='{.spec.nodeName}{"\n"}'
kubectl get pod <pod-name> -n <namespace> \
-o jsonpath='{range .spec.containers[*]}{.name}{" spec\t"}{.resources}{"\n"}{end}'
kubectl get pod <pod-name> -n <namespace> \
-o jsonpath='{range .status.containerStatuses[*]}{.name}{" status\t"}{.resources}{"\n"}{end}'
On that node, as root, with the systemd cgroup driver and cgroup v2 (these snippets also need jq). Pin the container with kubelet labels; a container name alone is not unique on a node. Use crictl ps -a if the container is not running:
CONTAINER_ID=$(crictl ps \
--label io.kubernetes.pod.namespace=<namespace> \
--label io.kubernetes.pod.name=<pod-name> \
--name <container-name> -q | head -n1)
# OCI view (not the CRI protobuf field names)
crictl inspect "$CONTAINER_ID" | jq '.info.runtimeSpec.linux.resources'
# Kernel cgroup path for this container
PID=$(crictl inspect "$CONTAINER_ID" | jq -r '.info.pid')
CGROUP="/sys/fs/cgroup$(awk -F: '$1=="0"{print $3}' /proc/$PID/cgroup)"
cat "$CGROUP/cpu.weight" # request.cpu → shares → weight; not a limit
cat "$CGROUP/cpu.max" # limit.cpu as "quota period"; unlimited is "max <period>"
cat "$CGROUP/memory.max" # limit.memory; unlimited is "max"
# Pod-level cgroup (parent slice), used by Pod-level resources
cat "$(dirname "$CGROUP")/cpu.max" "$(dirname "$CGROUP")/memory.max"
# Present when MemoryQoS is enabled
cat "$CGROUP/memory.high" "$CGROUP/memory.min" "$CGROUP/memory.low"
SCOPE=$(basename "$CGROUP")
systemctl show "$SCOPE" \
-p CPUWeight -p CPUQuotaPerSecUSec -p CPUQuotaPeriodUSec -p MemoryMax
Do not treat info.runtimeSpec.linux.cgroupsPath as a filesystem path when the runtime uses the systemd driver; that value is a systemd unit path (slice:runtime:id), not a directory under /sys/fs/cgroup.
Expected mapping at each layer:
- Kubernetes:
spec.containers[*].resources(desired) andstatus.containerStatuses[*].resources(enacted after in-place resize). Pod-level resources are onspec.resourcesand the parent pod slice. - CRI (kubelet → runtime; these protobuf names stay cgroup v1-style even on v2):
cpu_shares,cpu_quota,cpu_period,memory_limit_in_bytes, plusunifiedfor Memory QoS.crictl inspectdoes not show these names; it shows the OCI spec the runtime derived from them. - OCI spec:
linux.resources.cpu.{shares,quota,period},linux.resources.memory.limit, andlinux.resources.unified. The runtime converts shares, quota, and period into cgroup v2 files. After upgrading crun or runc, the shares-to-weight formula may change; see CPU weight conversion. - systemd scope:
CPUWeight,CPUQuotaPerSecUSec,CPUQuotaPeriodUSec,MemoryMax. Unlimited or unset values often appear as[not set]or infinity. - cgroupfs:
cpu.weight(from CPU request; an unset request still yields the default of 2 shares),cpu.max, andmemory.maxon the container scope. Memory QoS addsmemory.high,memory.min, andmemory.low.
Further reading
- Kubernetes 1.31: Moving cgroup v1 Support into Maintenance Mode
- Kubernetes v1.36: Tiered Memory Protection with Memory QoS
- Kubernetes v1.36: PSI Metrics for Kubernetes Graduates to GA
- Kubernetes 1.25: cgroup v2 graduates to GA
- KubeCon NA 2022 cgroup v2: Before You Jump In by Tony Gosselin & Mike Tougeron, Adobe Systems
- KubeCon NA 2022 Cgroupv2 Is Coming Soon To a Cluster Near You - David Porter, Google & Mrunal Patel, RedHat
- KubeCon EU 2020 Kubernetes On cgroup v2 by Giuseppe Scrivano, Red Hat.
- This blog only covers the basic requirements and configuration of Kubernetes components. It will not include how to enable cgroup fs in OS distributions. For migration, you can refer to migrating cgroup v2
06 Oct 2026 6:00pm GMT
05 Oct 2026
Kubernetes Blog
Scaling Kubernetes Workloads with Node Swap
Memory is often the first hard limit a Kubernetes cluster hits. Nodes run out of RAM long before they run out of CPU, and the new wave of agentic AI workloads makes this worse. These workloads demand large memory footprints to start up and run untrusted code, then sit idle waiting for the next prompt. That idle but resident memory is expensive, and it caps how many pods a node can hold. This is where swap helps. Kubernetes support for running nodes with swap enabled reached General Availability in v1.34, and by backing that swap with fast NVMe solid state drives (SSDs), a node can page out dormant memory and pack in far more pods. This post explains how we benchmarked that approach across three workloads, including CI/CD kernel builds, sandboxed headless browsers, and isolated Python runtimes; we found density gains of up to 3×, often with little or no latency cost.
The node density problem
The Kubernetes ecosystem has reached a fundamental physical resource constraint: the strict limits of hardware memory versus the growing demand for dynamic, bursty workloads in the new agentic era.
Historically, administrators provisioning memory-intensive workloads encountered a persistent dilemma: set memory limits too high and you waste expensive infrastructure on idle RAM; set them too low and you risk Out-Of-Memory (OOM) kills.
This conflict is amplified when deploying autonomous AI agents using secure execution environments like the agent-sandbox framework. These agentic pods require large memory footprints to initialize and execute untrusted code. However, after their burst of activity, they typically enter long-tail idle phases waiting for user prompts. Keeping this idle state in physical RAM caps cluster density and makes AI infrastructure expensive to run.
The solution: Kubernetes node swap
With the introduction of Kubernetes' support for running nodes with swap enabled (which reached General Availability in v1.34), this paradigm shifts. By enabling the Linux kernel to page out anonymous memory to disk, node swap acts as a shock absorber during traffic spikes or periods of heavy memory oversubscription.
Historically, swap was discouraged in Kubernetes for two reasons. The first was memory accounting. Under cgroup v1, the controls treated memory and swap as a single combined limit rather than letting operators set an independent limit for disk swap. Without independent tracking, a process could page large amounts of anonymous memory out to disk, which made a container's real memory usage unpredictable and hard to isolate. Kubernetes' swap support resolves this by relying on cgroup v2, whose separate swap accounting tracks disk swap on its own. The second reason was the latency penalty of paging to slow spinning disks, which fast NVMe Local SSDs largely eliminate. Together, these make it practical to increase pod density and buffer against memory spikes without sacrificing cluster stability.
The benchmark data
To quantify the performance boundaries and cost-saving potential of Local SSD-backed node swap, this analysis covers three distinct workload categories: a Traditional Build Workload for CI/CD pipelines, High-Density Browser Sandboxes, and Isolated Python Sandboxes.
| Workload Profile | Baseline Capacity (No Swap) | Local SSD Swap Capacity | Density Improvement |
|---|---|---|---|
| Linux CI/CD Kernel Build | 600 MB RAM Limit | 300 MB RAM Limit | -50% RAM Footprint |
| Headless Chrome (Kata) | 40 Concurrent Pods | 50 Concurrent Pods | +25% Pod Density |
| Headless Chrome (gVisor) | 80 Concurrent Pods | 160 Concurrent Pods | +100% Pod Density |
| Python Sandbox (gVisor) | 80 Concurrent Pods | 240 Concurrent Pods | +200% Pod Density |
1. Traditional workload: Linux kernel build
Before exploring specialized agentic architectures, swap was validated against classic batch workloads by running a complete Linux 6.1.1 kernel build. The kernel compilation process leverages concurrent worker threads, balloons in memory to hold compiled object files, and requires a large memory spike during the brief linking phase.
This workload mirrors the memory behavior of enterprise CI/CD pipelines. Because earlier compiled objects sit inactive in memory while the pipeline progresses, CI/CD jobs frequently hoard unused physical RAM, which makes them well suited to node swap compression.
On a baseline node without swap, the minimum memory limit to prevent an OOM crash during compilation was 600 MB. Routing swap to a Local SSD cut the container memory limit by 50% to 300 MB without incurring any execution slowdown (in fact, it ran cleanly in 374s vs the baseline 433s). However, as an explicit tradeoff, compressing the limit further to 200 MB forced the active working set into swap, causing long I/O wait times and increasing execution time by over 40%. This reinforces that swap serves as an insurance policy for burst memory, not a replacement for active RAM.
2. High-density agent workloads: headless browser runtimes
AI agent workloads frequently require manipulating headless browsers via Chromium. However, trusting external code execution often requires stricter security isolation than standard Linux namespaces. This benchmark cross-evaluated several container runtime environments. The raw logs and testing methodologies for the default runtime are available in the Agent Sandbox GKE Swap directory.
- Unsandboxed baseline limits (
runc): To test the limits of the environment without the overhead of security runtimes, plain runc containers were swept on a c4-standard-32 node (32 vCPU, 120 GB RAM). Without swap, the node exhausted physical memory and failed past 512 pods. Enabling Local SSD swap allowed the node to support 768 concurrent pods. - Advanced security runtimes (for example: gVisor, Kata Containers): Enabling strict security sandboxing increases memory overhead and normally reduces pod density. However, memory swap naturally absorbs this overhead penalty. Without swap, a gVisor environment hit a hard limit at 80 pods. Local SSD swap doubled that capacity, which allowed 160 concurrent gVisor pods on a single node. Similarly, Kata Containers microVMs exhausted physical RAM at 40 concurrent pods without swap, but using GCP Local SSD swap expanded this to 50 stable Kata microVMs before CPU saturation.
At these maximum densities, the per-pod latency increase is driven mainly by pods competing for CPU, not by swap I/O. An operator tuning for a specific latency target would run at a lower density than the peak numbers here and see a proportionally smaller latency cost.
For a comprehensive architectural breakdown and density metrics for gVisor and Kata, see the Agent Sandbox GKE Swap Runtimes directory.
3. Beyond browsers: sandboxed Python runtimes
The advantages of node swap also extend to untrusted, isolated code-execution environments. This sweep deployed simultaneous Python sandbox sessions analyzing 5 million rows of data from the MovieLens 20M dataset, requiring a ≃375 MiB resident memory footprint per execution. The in-depth scaling results and deployment code for this sweep can be reviewed in the Agent Sandbox GKE Swap Python Density directory.
Without swap, heavy concurrent bursts exhausted physical memory, causing the node to hit a hard RAM limit and fail at 80 concurrent sessions. Enabling Local SSD swap offloaded dormant anonymous memory, freeing up physical RAM and preserving the node's page cache. This allowed the node to scale to 240 concurrently isolated Python sandboxes-a 3× density improvement. As with the browser workloads, the latency rise at peak density comes mainly from the sandboxes competing for CPU rather than from swap itself.
Density Benchmarks with and without Node Swap
How to use it
If you manage Kubernetes infrastructure for developer environments, browser testing farms, JVM applications, or AI execution runtimes, leveraging Local SSD swap can multiply your density efficiency.
In Kubernetes v1.34+, node swap is Generally Available. You enable it via the kubelet configuration:
kind: KubeletConfiguration
apiVersion: kubelet.config.k8s.io/v1beta1
failSwapOn: false
memorySwap:
swapBehavior: LimitedSwap
Pairing this upstream configuration with your cloud provider's high-speed local disk gives you dynamic memory balancing. For example, this is natively supported on Google Kubernetes Engine via Node Memory Swap configured on Local SSD profiles.
To get the benefits, configure your workloads with Burstable QoS: set your container's memory limits higher than its requests. The node automatically rations fast swap space based on idle application memory usage while keeping active processes responsive.
Conclusion
As the Kubernetes ecosystem transitions into the agentic era, administrators face a growing conflict between finite hardware memory limits and the bursty behavior of AI workloads. Frameworks like Agent Sandbox provide the security isolation required for running untrusted agents, but that isolation traditionally demands large amounts of idle memory overhead.
By configuring the kubelet with LimitedSwap and routing it to Local SSDs, you can mitigate this conflict. Fast swap offloads the dormant states of idle agents, allowing you to increase pod density and node utilization on the same infrastructure without compromising security boundaries.
05 Oct 2026 6:00pm GMT
22 Sep 2026
Kubernetes Blog
Spotlight on SIG Apps
As Kubernetes adoption has grown, the conversation has shifted beyond running containers to managing increasingly complex application lifecycles. Modern platforms support stateless web services, stateful databases, batch processing, AI workloads, and platform services. At the same time, they must remain reliable during upgrades, scaling events, and infrastructure failures.
Every Kubernetes user relies on SIG Apps, whether they realize it or not. Deployments, StatefulSets, DaemonSets, Jobs, and CronJobs form the foundation of how applications are deployed, updated, scaled, and operated across the Kubernetes ecosystem.
SIG Apps is focused on improving workload resilience, refining application lifecycle management, and addressing the operational challenges that emerge when applications encounter node failures, rollout disruptions, and increasingly complex infrastructure environments.
In this spotlight, we sit down with two of the three SIG Apps chairs Janet Kuo and Maciej Szulik to discuss the evolution of Kubernetes workload management, the challenges of balancing application reliability with operational simplicity, and the future of application lifecycle management within one of Kubernetes' most influential Special Interest Groups.
Introducing SIG Apps
Natalie Fisher: Can you introduce yourself, your role, and how you got involved in SIG Apps?
Janet Kuo: I'm a Senior Staff Software Engineer at Google and have been a Kubernetes maintainer since 2015, joining the community just as we were racing toward the 1.0 launch. In those early days, my focus was on building the core Workloads API, specifically developing controllers like Deployment, ReplicaSet, StatefulSet, and DaemonSet, defining their rollout behaviors, and bringing them from initial designs to GA. That hands-on work was my entry point into SIG Apps.
Since then, I've stayed deeply involved in both the technical and community sides of Kubernetes. I have led SIG Apps as Co-Chair and Tech Lead since 2019. Currently, in addition to maintaining the workloads API, I am driving new subprojects like the Agent Sandbox to ensure Kubernetes is ready for next-generation agentic and AI workloads.
Maciej Szulik: I started contributing to Kubernetes all the way back in 2014. Since then, I've worked across various areas of the project: controllers, kubectl, and apimachinery, which eventually led me to become one of the Chairs and Tech Leads for SIG Apps. My current focus is reliability of the workload controllers under the SIG Apps umbrella and stability and ease of use of kubectl as part of my SIG CLI Tech Lead role. I also care about overall community health and growth as part of my Steering Committee role. Outside of Kubernetes, I work as a Staff Platform Engineer at Defense Unicorns, where I'm helping make Kubernetes more airgap-native with a project called zarf.
The problem and the solution
SIG Apps is responsible for the core workload APIs that power how applications run on Kubernetes. From Deployments and StatefulSets to Jobs and CronJobs, these controllers determine how workloads are created, updated, scaled, and recovered when things go wrong.
As Kubernetes expands to support increasingly diverse workloads - including AI, batch processing, and large-scale distributed applications - SIG Apps continues to evolve these APIs while balancing reliability, backward compatibility, and operational simplicity.
NF: For readers who may not be familiar, what is SIG Apps, and what role does it play within the broader Kubernetes ecosystem?
MS: SIG Apps is the Kubernetes Special Interest Group responsible for the workloads APIs. CronJob and Job help running batch workloads, whereas DaemonSet, Deployment, ReplicaSet, and StatefulSet serve the majority of other applications. More broadly, SIG Apps owns the layer most developers actually touch day-to-day: the controllers that turn a workload specification into running, self-healing pods. It's the group deciding how Deployments roll out, how Jobs retry, how DaemonSets place a pod per node.
JK: Adding to what Maciej described, as the industry shifts, we are seeing a massive demand to run complex, non-traditional workloads like distributed AI training, batch computing, and dynamic agent environments. Our role is expanding: we aren't just maintaining the classic workloads API, but we are actively evolving it and establishing new patterns (like the Agent Sandbox) to make sure Kubernetes remains the best platform for the next generation of workloads, such as AI.
NF: Looking at the workload APIs owned by SIG Apps (Deployments, StatefulSets, DaemonSets, Jobs, and CronJobs), which areas are receiving the most attention from maintainers and contributors?
MS: After a long stretch focused on making batch workloads run smoothly on Kubernetes, we've shifted attention to make sure serving workloads (DaemonSets, StatefulSets, etc) aren't left behind. This means performance and high-scale improvements to rollout and scaling behavior, plus working through our backlog of user-reported issues, prioritizing the ones with the strongest support from the user base.
Current focus areas
As Kubernetes workloads grow in scale and complexity, the challenges facing workload controllers evolve as well. We asked the SIG Apps chairs where contributors are focusing their efforts today and which resilience problems they believe are the highest priorities.
NF: From your perspective, what are the most important workload resilience problems SIG Apps is trying to solve today?
MS: Node lifecycle challenges have come up repeatedly across SIG Apps, SIG Node, and SIG Autoscaling discussions. DaemonSets and Jobs are just where the pain is most visible, since they're the workloads most directly bound to node state. Rather than solve it piecemeal within one SIG, we've settled on spinning up a dedicated Node Lifecycle Working Group to focus on this properly and hopefully land long-term solutions instead of one-off patches.
JK: From an AI perspective, resilience is critical. When you are running a massive distributed LLM training job that spans hundreds of GPUs, a single node failure can halt the entire pipeline. Similarly, if a DaemonSet that runs your logging or GPU monitoring agent gets stuck on a bad node, it impacts the entire cluster's health.
In addition to the work in the Node Lifecycle WG to handle infrastructure-level degradation, SIG Apps is addressing this at the orchestration layer through subprojects like JobSet (for distributed training) and LeaderWorkerSet (LWS) (for sharded LLM inference). These APIs introduce patterns like "all-or-nothing" failure handling, where a single pod or job failure triggers a coordinated group-level restart to resume from the last clean checkpoint, rather than letting stuck workloads hang in an inconsistent state.
Real-world impact
The work happening within SIG Apps extends far beyond controller implementations and API design. We wanted to understand what these improvements mean in practice for platform teams operating Kubernetes clusters in production.
NF: For platform teams operating Kubernetes in production, what practical improvements would they notice if the node lifecycle and workload resilience work currently under discussion is successfully delivered?
MS: I'm mostly looking from the sidelines, the folks actually in the Node Lifecycle Working Group would give you a sharper answer. But from where I sit, I'm hoping their work translates into fewer 3am pages that turn out to be "a DaemonSet rollout got stuck because node X was flaky, and someone had to manually cordon/delete/restart to unstick it."
JK: +1 to what Maciej said, and beyond reducing manual intervention, platform teams will also see much better resource predictability and cost efficiency. For example, in AI workloads where GPU idle time is extremely expensive, having Kubernetes automatically detect a degraded node and reschedule the training coordinator or agent before the job crashes means less wasted compute and more stable job execution.
Challenges and trade-offs
Evolving APIs that millions of workloads rely on requires careful engineering and even more careful decision-making. We asked the SIG Apps chairs about the technical and operational trade-offs they weigh when introducing changes to Kubernetes' core workload controllers.
NF: What are some of the hardest technical or operational trade-offs SIG Apps encounters when evolving core workload controllers?
MS: Honestly, a few tensions keep coming up: how aggressively a controller should give up on stuck pods, and what signals it actually needs to make that call correctly. At the same time, we always have to think about backward compatibility. Deployment, DaemonSet, and Job behavior has been depended on for a decade [by Kubernetes users, tooling, automation, and higher-level controllers], so even a change that's clearly "more correct" can break automation people built around the old behavior without meaning to.
JK: One of our hardest trade-offs is resisting the urge to make "elegant" design changes that break backward compatibility. Instead, we have to design opt-in features that let users adopt new behaviors without forcing them on legacy workloads. When we need to support completely new paradigms, we prefer introducing them as CRDs first rather than bloating the core APIs, like we are doing with Agent Sandbox, JobSet, and LWS.
Looking ahead
While much of SIG Apps' work focuses on maintaining the stability of existing workload APIs, the group is also shaping the future of Kubernetes through new enhancements and proposals. We concluded by asking about one proposal that recently returned to active development and what it represents for the future of workload management.
NF: The SIG recently discussed reviving KEP-4443 with a target release of Kubernetes 1.38. What opportunities or challenges does this proposal aim to address, and why is now the right time to revisit it?
KEP-4443 addresses a small but real gap in the Job API: a PodFailurePolicy can be configured to add a condition reason to the JobFailed condition, but different pod failure policy rules targeting different container exit codes all produce that same generic reason. The proposal is simple: an optional Name field on each PodFailurePolicyRule, which gets appended to the JobFailed condition reason, so higher-level tools like JobSet can finally react differently depending on which rule triggered the failure.
As for timing, the answer is as simple as it always is in open source: we lost the original contributor who was driving this. Now we've got someone new interested in picking it up, that's why we're targeting the next release.
Getting Involved
NF: For someone interested in contributing to SIG Apps, where would you recommend they start, especially if they are not yet a Kubernetes maintainer?
MS: The best place to start is the #sig-apps slack channel and our regular SIG Apps meetings. We've all started there, and if it feels intimidating, or nobody replies right away, that's completely normal. Everyone's busy. It's not personal.
JK: In addition to what Maciej answered, I'd suggest looking at our newer subprojects and initiatives. Contributing to stable APIs like Deployment or StatefulSet can be daunting because the barrier for making changes is very high due to backward compatibility, and there is much less low-hanging fruit.
If you are new to the community, projects like the Agent Sandbox are fantastic entry points. They are actively evolving, have a friendly group of maintainers, and offer plenty of greenfield development opportunities where you can make a significant impact quickly.
Summary
SIG Apps has shaped how Kubernetes applications are deployed and operated since the project's earliest days. While users often interact with Deployments, StatefulSets, Jobs, and DaemonSets without thinking about the controllers behind them, the work within SIG Apps continues to shape the reliability and scalability of workloads across the Kubernetes ecosystem.
From improving workload resilience and node lifecycle behavior to enabling new patterns for AI and distributed computing, the SIG is evolving Kubernetes while remaining committed to one of the project's core principles: preserving the stability and backward compatibility that users depend on. Whether you're interested in core workload APIs, emerging projects like Agent Sandbox, or helping improve the operational experience of Kubernetes users everywhere, SIG Apps offers many opportunities to get involved.
22 Sep 2026 6:00pm GMT