06 Oct 2026

feedKubernetes Blog

The Shift to cgroup v2 in Kubernetes: What You Need to Know

In Linux, cgroups (control groups) are a kernel feature used for managing system resources. Kubernetes uses cgroups to allocate resources like CPU and memory to containers, ensuring that applications run smoothly without interfering with each other. With the release of Kubernetes v1.31, support for v1 cgroup management moved into maintenance mode. Support for v2 cgroup management has been stable since Kubernetes v1.25.

Compared with cgroup v1, cgroup v2 provides a single unified hierarchy, a more consistent interface, and a stronger foundation for resource isolation and modern resource-management features.

Deprecation of cgroup v1

Kubernetes has deprecated cgroup v1. Starting with Kubernetes v1.35, failCgroupV1 defaults to true, so the kubelet does not start on a cgroup v1 node by default. Administrators can temporarily set failCgroupV1: false in the kubelet configuration file, but removal will follow the Kubernetes deprecation policy. Further removal work is tracked in KEP-5573: Remove cgroup v1 support.

If you are still on a release older than v1.35, migrate every Linux node to cgroup v2 before upgrading, or plan to set the temporary failCgroupV1: false override. If you are already on v1.35 or later, confirm that every Linux node runs cgroup v2 (or that you intentionally keep the override). Under the default configuration, a remaining cgroup v1 node fails during kubelet startup.

For kubeadm-managed clusters, Kubernetes v1.35 also makes this an earlier, stricter check. The SystemVerification preflight check, provided by k8s.io/system-validators, returns an error during kubeadm init, kubeadm join, and kubeadm upgrade when it detects cgroup v1 with kubelet v1.35 or later; with an older kubelet, the check remains a warning. See kubernetes/system-validators#1.12.1 release notes for details.

The top FAQs cover three main areas: why to migrate, the benefits and drawbacks, and key points to keep in mind when using cgroup v2.

Limitations of cgroup v1 and Improvements with cgroup v2

The Linux kernel documentation describes both interfaces:

Let's enumerate some known issues.

active_file memory is not considered available memory

The kubelet treats active_file memory as not reclaimable. For I/O-intensive workloads, a large page cache can therefore make the kubelet report memory pressure and evict Pods. This is a known kubelet issue (kubernetes/kubernetes#43916); migrating to cgroup v2 does not by itself change that calculation. The documented workaround is to set equal memory requests and limits for containers that perform intensive I/O, after measuring an appropriate value.

Memory QoS updates in Kubernetes v1.36

Memory QoS was introduced as an alpha feature in Kubernetes v1.22 and updated in v1.27. It remains alpha in v1.36, but now separates memory throttling from memory reservation and adds tiered memory protection:

Memory QoS is available only on Linux nodes that use cgroup v2. It relies on the cgroup v2 memory controller: memory.high provides throttling, while memory.min and memory.low provide hard and soft protection when tiered reservation is enabled. cgroup v1 cannot provide this protection model.

For example, to opt in to tiered protection:

apiVersion: kubelet.config.k8s.io/v1beta1
kind: KubeletConfiguration
featureGates:
 MemoryQoS: true
memoryReservationPolicy: TieredReservation
memoryThrottlingFactor: 0.9

The overall Kubernetes recommendation is not to enable Alpha features in production; however, if you judge that memory QoS with tiered reservations is useful for your platform in production, make sure to test the configuration and account for hard-reserved memory before you enable the TieredReservation feature gate. See the Pod QoS documentation for the current mapping of Kubernetes QoS classes to cgroup v2 controls.

Container-aware OOM handling

On cgroup v2 nodes, the kubelet defaults singleProcessOOMKill to false. It therefore sets memory.oom.group for each container cgroup so that an OOM event kills all processes in that container together, rather than leaving a partially functioning multi-process container. Set singleProcessOOMKill: true only if you need the cgroup v1-compatible behavior where the kernel may kill one process at a time. See the KubeletConfiguration reference for this setting.

This behavior is scoped to a container cgroup, not the whole Pod. Also, cgroup.kill is a separate administrative interface: writing 1 to it sends SIGKILL to every process in that cgroup and its descendants; it does not configure OOM behavior. The cgroup v2 memory controller additionally provides memory.events counters that monitoring systems and userspace OOM managers can observe.

Rootless support

In cgroup v1, delegating controllers to less privileged containers may be dangerous.

Unlike cgroup v1, cgroup v2 officially supports delegation. Most implementations of rootless containers rely on systemd for delegating v2 controllers to non-root users.

This delegation mechanism is separate from Kubernetes Pod user namespaces, which map container users to unprivileged host users. Pod user namespace support graduated to stable in Kubernetes v1.36; check its filesystem, kernel, CRI runtime, and OCI runtime prerequisites before enabling it.

What else?

  1. eBPF stories:
    • In cgroup v1, device access controls are exposed through interface files.
    • The cgroup v2 device controller has no interface files, and is implemented on top of cgroup BPF.
    • Cilium attaches BPF cgroup programs for socket-based load balancing. Its default cgroup root is /run/cilium/cgroupv2.
  2. Pressure Stall Information (PSI) reports CPU, memory, and I/O contention at node, Pod, and container level. On supported clusters, the kubelet exposes PSI by default (KubeletPSI is stable and locked on). PSI requires cgroup v2, Linux 4.20 or later, CONFIG_PSI=y, and a kernel not booted with psi=0. The kubelet surfaces the data through the Summary API and /metrics/cadvisor.
  3. When migrating, update software that reads the cgroup filesystem directly. The migration guidance recommends cAdvisor v0.43.0 or later and lists compatible Java, Node.js, and automaxprocs versions.

CPU weight conversion in newer OCI runtimes

cgroup v1 uses cpu.shares, whereas cgroup v2 uses cpu.weight. Newer OCI runtimes use an improved non-linear conversion that preserves the default priority and gives small CPU requests more usable granularity. The change is implemented in the OCI runtime rather than Kubernetes: it is available in crun v1.23 and runc v1.3.2. After upgrading a runtime, monitoring or policy tools that predict exact cpu.weight values may need updates. Read New Conversion from cgroup v1 CPU Shares to v2 CPU Weight for the formula, examples, and compatibility considerations.

In-place resource updates

In-place Pod vertical scaling graduated to stable in Kubernetes v1.35. Kubernetes v1.36 then [enabled] (/blog/2026/04/30/kubernetes-v1-36-inplace-pod-level-resources-beta/) in-place vertical scaling for Pod-level resources by default, as a Beta feature. The kubelet coordinates changes between the Pod-level and container cgroups so that increases create headroom before container limits grow, while decreases constrain containers before shrinking the Pod-level boundary. Accurate aggregate enforcement for this v1.36 feature requires cgroup v2.

Adopting cgroup version 2

Requirements

Here's what you need to use cgroup v2 with Kubernetes. First up, you need to be using a version of Kubernetes with support for v2 cgroup management; that's been stable since Kubernetes v1.25 and all supported Kubernetes releases include this support.

For now, you can opt back in to use cgroup v1; the Kubernetes project recommends using cgroup v2, but in Kubernetes 1.36 (the current release) the cgroup v1 option remains supported as a fallback. That fallback is scheduled for removal in Kubernetes v1.38. If you are running an older cluster, plan to migrate; if you are setting up a new cluster with Linux nodes, you should prefer cgroup v2. In either case, review both the Kubernetes runtime documentation and (if relevant) the containerd compatibility matrix.

kernel updates around cgroup v2

When Kubernetes was first announced, in 2014, only v1 cgroup existed. Version 2 cgroup management first appeared in Linux kernel 4.5, released in 2016.

cgroup driver configuration

Configure the kubelet's cgroup driver to match the container runtime cgroup driver.

If you use kubeadm to manage your cluster, Kubernetes recommends that you use the systemd cgroup driver, because kubeadm manages the kubelet as a systemd service. For other management tooling, check the documentation for the tool you're using to manage your cluster.

If you can pick either option, I recommend using the systemd driver.

Whatever tooling you've chosen, the kubelet automatically tries to detect the runtime's recommended cgroup driver. This automatic detection relies on using a runtime that implements the RuntimeConfig CRI RPC (for example: containerd v2.0+ or CRI-O v1.28+).

If you're using a container runtime that supports cgroup v2 but doesn't support automatic cgroup driver detection, you can manually configure an override by editing the kubelet configuration file. For example:

apiVersion: kubelet.config.k8s.io/v1beta1
kind: KubeletConfiguration
cgroupDriver: systemd

Automatic discovery of the runtime's cgroup driver through the CRI, tracked by KEP-4033, graduated to stable in Kubernetes v1.34. It requires a runtime that implements the RuntimeConfig CRI RPC (containerd v2.0+ or CRI-O v1.28+). When available, the kubelet uses the value reported by the runtime instead of its configured cgroupDriver value.

Tools and commands for troubleshooting

Tools and commands that you should know about cgroups:

How to check if a Pod CPU or memory limit is successfully applied to the cgroup file?

Work from the API object down to the node. Identify the node, then compare desired resources in the spec with the enacted values in status (in-place resize can leave those out of sync):

kubectl get pod <pod-name> -n <namespace> \
 -o jsonpath='{.spec.nodeName}{"\n"}'

kubectl get pod <pod-name> -n <namespace> \
 -o jsonpath='{range .spec.containers[*]}{.name}{" spec\t"}{.resources}{"\n"}{end}'

kubectl get pod <pod-name> -n <namespace> \
 -o jsonpath='{range .status.containerStatuses[*]}{.name}{" status\t"}{.resources}{"\n"}{end}'

On that node, as root, with the systemd cgroup driver and cgroup v2 (these snippets also need jq). Pin the container with kubelet labels; a container name alone is not unique on a node. Use crictl ps -a if the container is not running:

CONTAINER_ID=$(crictl ps \
 --label io.kubernetes.pod.namespace=<namespace> \
 --label io.kubernetes.pod.name=<pod-name> \
 --name <container-name> -q | head -n1)

# OCI view (not the CRI protobuf field names)
crictl inspect "$CONTAINER_ID" | jq '.info.runtimeSpec.linux.resources'

# Kernel cgroup path for this container
PID=$(crictl inspect "$CONTAINER_ID" | jq -r '.info.pid')
CGROUP="/sys/fs/cgroup$(awk -F: '$1=="0"{print $3}' /proc/$PID/cgroup)"

cat "$CGROUP/cpu.weight" # request.cpu → shares → weight; not a limit
cat "$CGROUP/cpu.max" # limit.cpu as "quota period"; unlimited is "max <period>"
cat "$CGROUP/memory.max" # limit.memory; unlimited is "max"

# Pod-level cgroup (parent slice), used by Pod-level resources
cat "$(dirname "$CGROUP")/cpu.max" "$(dirname "$CGROUP")/memory.max"

# Present when MemoryQoS is enabled
cat "$CGROUP/memory.high" "$CGROUP/memory.min" "$CGROUP/memory.low"

SCOPE=$(basename "$CGROUP")
systemctl show "$SCOPE" \
 -p CPUWeight -p CPUQuotaPerSecUSec -p CPUQuotaPeriodUSec -p MemoryMax

Do not treat info.runtimeSpec.linux.cgroupsPath as a filesystem path when the runtime uses the systemd driver; that value is a systemd unit path (slice:runtime:id), not a directory under /sys/fs/cgroup.

Expected mapping at each layer:

Further reading

06 Oct 2026 6:00pm GMT

05 Oct 2026

feedKubernetes Blog

Scaling Kubernetes Workloads with Node Swap

Memory is often the first hard limit a Kubernetes cluster hits. Nodes run out of RAM long before they run out of CPU, and the new wave of agentic AI workloads makes this worse. These workloads demand large memory footprints to start up and run untrusted code, then sit idle waiting for the next prompt. That idle but resident memory is expensive, and it caps how many pods a node can hold. This is where swap helps. Kubernetes support for running nodes with swap enabled reached General Availability in v1.34, and by backing that swap with fast NVMe solid state drives (SSDs), a node can page out dormant memory and pack in far more pods. This post explains how we benchmarked that approach across three workloads, including CI/CD kernel builds, sandboxed headless browsers, and isolated Python runtimes; we found density gains of up to 3×, often with little or no latency cost.

The node density problem

The Kubernetes ecosystem has reached a fundamental physical resource constraint: the strict limits of hardware memory versus the growing demand for dynamic, bursty workloads in the new agentic era.

Historically, administrators provisioning memory-intensive workloads encountered a persistent dilemma: set memory limits too high and you waste expensive infrastructure on idle RAM; set them too low and you risk Out-Of-Memory (OOM) kills.

This conflict is amplified when deploying autonomous AI agents using secure execution environments like the agent-sandbox framework. These agentic pods require large memory footprints to initialize and execute untrusted code. However, after their burst of activity, they typically enter long-tail idle phases waiting for user prompts. Keeping this idle state in physical RAM caps cluster density and makes AI infrastructure expensive to run.

The solution: Kubernetes node swap

With the introduction of Kubernetes' support for running nodes with swap enabled (which reached General Availability in v1.34), this paradigm shifts. By enabling the Linux kernel to page out anonymous memory to disk, node swap acts as a shock absorber during traffic spikes or periods of heavy memory oversubscription.

Historically, swap was discouraged in Kubernetes for two reasons. The first was memory accounting. Under cgroup v1, the controls treated memory and swap as a single combined limit rather than letting operators set an independent limit for disk swap. Without independent tracking, a process could page large amounts of anonymous memory out to disk, which made a container's real memory usage unpredictable and hard to isolate. Kubernetes' swap support resolves this by relying on cgroup v2, whose separate swap accounting tracks disk swap on its own. The second reason was the latency penalty of paging to slow spinning disks, which fast NVMe Local SSDs largely eliminate. Together, these make it practical to increase pod density and buffer against memory spikes without sacrificing cluster stability.

The benchmark data

To quantify the performance boundaries and cost-saving potential of Local SSD-backed node swap, this analysis covers three distinct workload categories: a Traditional Build Workload for CI/CD pipelines, High-Density Browser Sandboxes, and Isolated Python Sandboxes.

Workload Profile Baseline Capacity (No Swap) Local SSD Swap Capacity Density Improvement
Linux CI/CD Kernel Build 600 MB RAM Limit 300 MB RAM Limit -50% RAM Footprint
Headless Chrome (Kata) 40 Concurrent Pods 50 Concurrent Pods +25% Pod Density
Headless Chrome (gVisor) 80 Concurrent Pods 160 Concurrent Pods +100% Pod Density
Python Sandbox (gVisor) 80 Concurrent Pods 240 Concurrent Pods +200% Pod Density

1. Traditional workload: Linux kernel build

Before exploring specialized agentic architectures, swap was validated against classic batch workloads by running a complete Linux 6.1.1 kernel build. The kernel compilation process leverages concurrent worker threads, balloons in memory to hold compiled object files, and requires a large memory spike during the brief linking phase.

This workload mirrors the memory behavior of enterprise CI/CD pipelines. Because earlier compiled objects sit inactive in memory while the pipeline progresses, CI/CD jobs frequently hoard unused physical RAM, which makes them well suited to node swap compression.

On a baseline node without swap, the minimum memory limit to prevent an OOM crash during compilation was 600 MB. Routing swap to a Local SSD cut the container memory limit by 50% to 300 MB without incurring any execution slowdown (in fact, it ran cleanly in 374s vs the baseline 433s). However, as an explicit tradeoff, compressing the limit further to 200 MB forced the active working set into swap, causing long I/O wait times and increasing execution time by over 40%. This reinforces that swap serves as an insurance policy for burst memory, not a replacement for active RAM.

2. High-density agent workloads: headless browser runtimes

AI agent workloads frequently require manipulating headless browsers via Chromium. However, trusting external code execution often requires stricter security isolation than standard Linux namespaces. This benchmark cross-evaluated several container runtime environments. The raw logs and testing methodologies for the default runtime are available in the Agent Sandbox GKE Swap directory.

At these maximum densities, the per-pod latency increase is driven mainly by pods competing for CPU, not by swap I/O. An operator tuning for a specific latency target would run at a lower density than the peak numbers here and see a proportionally smaller latency cost.

For a comprehensive architectural breakdown and density metrics for gVisor and Kata, see the Agent Sandbox GKE Swap Runtimes directory.

3. Beyond browsers: sandboxed Python runtimes

The advantages of node swap also extend to untrusted, isolated code-execution environments. This sweep deployed simultaneous Python sandbox sessions analyzing 5 million rows of data from the MovieLens 20M dataset, requiring a ≃375 MiB resident memory footprint per execution. The in-depth scaling results and deployment code for this sweep can be reviewed in the Agent Sandbox GKE Swap Python Density directory.

Without swap, heavy concurrent bursts exhausted physical memory, causing the node to hit a hard RAM limit and fail at 80 concurrent sessions. Enabling Local SSD swap offloaded dormant anonymous memory, freeing up physical RAM and preserving the node's page cache. This allowed the node to scale to 240 concurrently isolated Python sandboxes-a 3× density improvement. As with the browser workloads, the latency rise at peak density comes mainly from the sandboxes competing for CPU rather than from swap itself.

Density Benchmarks with and without Node Swap

How to use it

If you manage Kubernetes infrastructure for developer environments, browser testing farms, JVM applications, or AI execution runtimes, leveraging Local SSD swap can multiply your density efficiency.

In Kubernetes v1.34+, node swap is Generally Available. You enable it via the kubelet configuration:

kind: KubeletConfiguration
apiVersion: kubelet.config.k8s.io/v1beta1
failSwapOn: false
memorySwap:
 swapBehavior: LimitedSwap

Pairing this upstream configuration with your cloud provider's high-speed local disk gives you dynamic memory balancing. For example, this is natively supported on Google Kubernetes Engine via Node Memory Swap configured on Local SSD profiles.

To get the benefits, configure your workloads with Burstable QoS: set your container's memory limits higher than its requests. The node automatically rations fast swap space based on idle application memory usage while keeping active processes responsive.

Conclusion

As the Kubernetes ecosystem transitions into the agentic era, administrators face a growing conflict between finite hardware memory limits and the bursty behavior of AI workloads. Frameworks like Agent Sandbox provide the security isolation required for running untrusted agents, but that isolation traditionally demands large amounts of idle memory overhead.

By configuring the kubelet with LimitedSwap and routing it to Local SSDs, you can mitigate this conflict. Fast swap offloads the dormant states of idle agents, allowing you to increase pod density and node utilization on the same infrastructure without compromising security boundaries.

05 Oct 2026 6:00pm GMT

22 Sep 2026

feedKubernetes Blog

Spotlight on SIG Apps

As Kubernetes adoption has grown, the conversation has shifted beyond running containers to managing increasingly complex application lifecycles. Modern platforms support stateless web services, stateful databases, batch processing, AI workloads, and platform services. At the same time, they must remain reliable during upgrades, scaling events, and infrastructure failures.

Every Kubernetes user relies on SIG Apps, whether they realize it or not. Deployments, StatefulSets, DaemonSets, Jobs, and CronJobs form the foundation of how applications are deployed, updated, scaled, and operated across the Kubernetes ecosystem.

SIG Apps is focused on improving workload resilience, refining application lifecycle management, and addressing the operational challenges that emerge when applications encounter node failures, rollout disruptions, and increasingly complex infrastructure environments.

In this spotlight, we sit down with two of the three SIG Apps chairs Janet Kuo and Maciej Szulik to discuss the evolution of Kubernetes workload management, the challenges of balancing application reliability with operational simplicity, and the future of application lifecycle management within one of Kubernetes' most influential Special Interest Groups.

Introducing SIG Apps

Natalie Fisher: Can you introduce yourself, your role, and how you got involved in SIG Apps?

Janet Kuo: I'm a Senior Staff Software Engineer at Google and have been a Kubernetes maintainer since 2015, joining the community just as we were racing toward the 1.0 launch. In those early days, my focus was on building the core Workloads API, specifically developing controllers like Deployment, ReplicaSet, StatefulSet, and DaemonSet, defining their rollout behaviors, and bringing them from initial designs to GA. That hands-on work was my entry point into SIG Apps.

Since then, I've stayed deeply involved in both the technical and community sides of Kubernetes. I have led SIG Apps as Co-Chair and Tech Lead since 2019. Currently, in addition to maintaining the workloads API, I am driving new subprojects like the Agent Sandbox to ensure Kubernetes is ready for next-generation agentic and AI workloads.

Maciej Szulik: I started contributing to Kubernetes all the way back in 2014. Since then, I've worked across various areas of the project: controllers, kubectl, and apimachinery, which eventually led me to become one of the Chairs and Tech Leads for SIG Apps. My current focus is reliability of the workload controllers under the SIG Apps umbrella and stability and ease of use of kubectl as part of my SIG CLI Tech Lead role. I also care about overall community health and growth as part of my Steering Committee role. Outside of Kubernetes, I work as a Staff Platform Engineer at Defense Unicorns, where I'm helping make Kubernetes more airgap-native with a project called zarf.

The problem and the solution

SIG Apps is responsible for the core workload APIs that power how applications run on Kubernetes. From Deployments and StatefulSets to Jobs and CronJobs, these controllers determine how workloads are created, updated, scaled, and recovered when things go wrong.

As Kubernetes expands to support increasingly diverse workloads - including AI, batch processing, and large-scale distributed applications - SIG Apps continues to evolve these APIs while balancing reliability, backward compatibility, and operational simplicity.

NF: For readers who may not be familiar, what is SIG Apps, and what role does it play within the broader Kubernetes ecosystem?

MS: SIG Apps is the Kubernetes Special Interest Group responsible for the workloads APIs. CronJob and Job help running batch workloads, whereas DaemonSet, Deployment, ReplicaSet, and StatefulSet serve the majority of other applications. More broadly, SIG Apps owns the layer most developers actually touch day-to-day: the controllers that turn a workload specification into running, self-healing pods. It's the group deciding how Deployments roll out, how Jobs retry, how DaemonSets place a pod per node.

JK: Adding to what Maciej described, as the industry shifts, we are seeing a massive demand to run complex, non-traditional workloads like distributed AI training, batch computing, and dynamic agent environments. Our role is expanding: we aren't just maintaining the classic workloads API, but we are actively evolving it and establishing new patterns (like the Agent Sandbox) to make sure Kubernetes remains the best platform for the next generation of workloads, such as AI.

NF: Looking at the workload APIs owned by SIG Apps (Deployments, StatefulSets, DaemonSets, Jobs, and CronJobs), which areas are receiving the most attention from maintainers and contributors?

MS: After a long stretch focused on making batch workloads run smoothly on Kubernetes, we've shifted attention to make sure serving workloads (DaemonSets, StatefulSets, etc) aren't left behind. This means performance and high-scale improvements to rollout and scaling behavior, plus working through our backlog of user-reported issues, prioritizing the ones with the strongest support from the user base.

Current focus areas

As Kubernetes workloads grow in scale and complexity, the challenges facing workload controllers evolve as well. We asked the SIG Apps chairs where contributors are focusing their efforts today and which resilience problems they believe are the highest priorities.

NF: From your perspective, what are the most important workload resilience problems SIG Apps is trying to solve today?

MS: Node lifecycle challenges have come up repeatedly across SIG Apps, SIG Node, and SIG Autoscaling discussions. DaemonSets and Jobs are just where the pain is most visible, since they're the workloads most directly bound to node state. Rather than solve it piecemeal within one SIG, we've settled on spinning up a dedicated Node Lifecycle Working Group to focus on this properly and hopefully land long-term solutions instead of one-off patches.

JK: From an AI perspective, resilience is critical. When you are running a massive distributed LLM training job that spans hundreds of GPUs, a single node failure can halt the entire pipeline. Similarly, if a DaemonSet that runs your logging or GPU monitoring agent gets stuck on a bad node, it impacts the entire cluster's health.

In addition to the work in the Node Lifecycle WG to handle infrastructure-level degradation, SIG Apps is addressing this at the orchestration layer through subprojects like JobSet (for distributed training) and LeaderWorkerSet (LWS) (for sharded LLM inference). These APIs introduce patterns like "all-or-nothing" failure handling, where a single pod or job failure triggers a coordinated group-level restart to resume from the last clean checkpoint, rather than letting stuck workloads hang in an inconsistent state.

Real-world impact

The work happening within SIG Apps extends far beyond controller implementations and API design. We wanted to understand what these improvements mean in practice for platform teams operating Kubernetes clusters in production.

NF: For platform teams operating Kubernetes in production, what practical improvements would they notice if the node lifecycle and workload resilience work currently under discussion is successfully delivered?

MS: I'm mostly looking from the sidelines, the folks actually in the Node Lifecycle Working Group would give you a sharper answer. But from where I sit, I'm hoping their work translates into fewer 3am pages that turn out to be "a DaemonSet rollout got stuck because node X was flaky, and someone had to manually cordon/delete/restart to unstick it."

JK: +1 to what Maciej said, and beyond reducing manual intervention, platform teams will also see much better resource predictability and cost efficiency. For example, in AI workloads where GPU idle time is extremely expensive, having Kubernetes automatically detect a degraded node and reschedule the training coordinator or agent before the job crashes means less wasted compute and more stable job execution.

Challenges and trade-offs

Evolving APIs that millions of workloads rely on requires careful engineering and even more careful decision-making. We asked the SIG Apps chairs about the technical and operational trade-offs they weigh when introducing changes to Kubernetes' core workload controllers.

NF: What are some of the hardest technical or operational trade-offs SIG Apps encounters when evolving core workload controllers?

MS: Honestly, a few tensions keep coming up: how aggressively a controller should give up on stuck pods, and what signals it actually needs to make that call correctly. At the same time, we always have to think about backward compatibility. Deployment, DaemonSet, and Job behavior has been depended on for a decade [by Kubernetes users, tooling, automation, and higher-level controllers], so even a change that's clearly "more correct" can break automation people built around the old behavior without meaning to.

JK: One of our hardest trade-offs is resisting the urge to make "elegant" design changes that break backward compatibility. Instead, we have to design opt-in features that let users adopt new behaviors without forcing them on legacy workloads. When we need to support completely new paradigms, we prefer introducing them as CRDs first rather than bloating the core APIs, like we are doing with Agent Sandbox, JobSet, and LWS.

Looking ahead

While much of SIG Apps' work focuses on maintaining the stability of existing workload APIs, the group is also shaping the future of Kubernetes through new enhancements and proposals. We concluded by asking about one proposal that recently returned to active development and what it represents for the future of workload management.

NF: The SIG recently discussed reviving KEP-4443 with a target release of Kubernetes 1.38. What opportunities or challenges does this proposal aim to address, and why is now the right time to revisit it?

KEP-4443 addresses a small but real gap in the Job API: a PodFailurePolicy can be configured to add a condition reason to the JobFailed condition, but different pod failure policy rules targeting different container exit codes all produce that same generic reason. The proposal is simple: an optional Name field on each PodFailurePolicyRule, which gets appended to the JobFailed condition reason, so higher-level tools like JobSet can finally react differently depending on which rule triggered the failure.

As for timing, the answer is as simple as it always is in open source: we lost the original contributor who was driving this. Now we've got someone new interested in picking it up, that's why we're targeting the next release.

Getting Involved

NF: For someone interested in contributing to SIG Apps, where would you recommend they start, especially if they are not yet a Kubernetes maintainer?

MS: The best place to start is the #sig-apps slack channel and our regular SIG Apps meetings. We've all started there, and if it feels intimidating, or nobody replies right away, that's completely normal. Everyone's busy. It's not personal.

JK: In addition to what Maciej answered, I'd suggest looking at our newer subprojects and initiatives. Contributing to stable APIs like Deployment or StatefulSet can be daunting because the barrier for making changes is very high due to backward compatibility, and there is much less low-hanging fruit.

If you are new to the community, projects like the Agent Sandbox are fantastic entry points. They are actively evolving, have a friendly group of maintainers, and offer plenty of greenfield development opportunities where you can make a significant impact quickly.

Summary

SIG Apps has shaped how Kubernetes applications are deployed and operated since the project's earliest days. While users often interact with Deployments, StatefulSets, Jobs, and DaemonSets without thinking about the controllers behind them, the work within SIG Apps continues to shape the reliability and scalability of workloads across the Kubernetes ecosystem.

From improving workload resilience and node lifecycle behavior to enabling new patterns for AI and distributed computing, the SIG is evolving Kubernetes while remaining committed to one of the project's core principles: preserving the stability and backward compatibility that users depend on. Whether you're interested in core workload APIs, emerging projects like Agent Sandbox, or helping improve the operational experience of Kubernetes users everywhere, SIG Apps offers many opportunities to get involved.

22 Sep 2026 6:00pm GMT