05 Oct 2026
Kubernetes Blog
Scaling Kubernetes Workloads with Node Swap
Memory is often the first hard limit a Kubernetes cluster hits. Nodes run out of RAM long before they run out of CPU, and the new wave of agentic AI workloads makes this worse. These workloads demand large memory footprints to start up and run untrusted code, then sit idle waiting for the next prompt. That idle but resident memory is expensive, and it caps how many pods a node can hold. This is where swap helps. Kubernetes support for running nodes with swap enabled reached General Availability in v1.34, and by backing that swap with fast NVMe solid state drives (SSDs), a node can page out dormant memory and pack in far more pods. This post explains how we benchmarked that approach across three workloads, including CI/CD kernel builds, sandboxed headless browsers, and isolated Python runtimes; we found density gains of up to 3×, often with little or no latency cost.
The node density problem
The Kubernetes ecosystem has reached a fundamental physical resource constraint: the strict limits of hardware memory versus the growing demand for dynamic, bursty workloads in the new agentic era.
Historically, administrators provisioning memory-intensive workloads encountered a persistent dilemma: set memory limits too high and you waste expensive infrastructure on idle RAM; set them too low and you risk Out-Of-Memory (OOM) kills.
This conflict is amplified when deploying autonomous AI agents using secure execution environments like the agent-sandbox framework. These agentic pods require large memory footprints to initialize and execute untrusted code. However, after their burst of activity, they typically enter long-tail idle phases waiting for user prompts. Keeping this idle state in physical RAM caps cluster density and makes AI infrastructure expensive to run.
The solution: Kubernetes node swap
With the introduction of Kubernetes' support for running nodes with swap enabled (which reached General Availability in v1.34), this paradigm shifts. By enabling the Linux kernel to page out anonymous memory to disk, node swap acts as a shock absorber during traffic spikes or periods of heavy memory oversubscription.
Historically, swap was discouraged in Kubernetes for two reasons. The first was memory accounting. Under cgroup v1, the controls treated memory and swap as a single combined limit rather than letting operators set an independent limit for disk swap. Without independent tracking, a process could page large amounts of anonymous memory out to disk, which made a container's real memory usage unpredictable and hard to isolate. Kubernetes' swap support resolves this by relying on cgroup v2, whose separate swap accounting tracks disk swap on its own. The second reason was the latency penalty of paging to slow spinning disks, which fast NVMe Local SSDs largely eliminate. Together, these make it practical to increase pod density and buffer against memory spikes without sacrificing cluster stability.
The benchmark data
To quantify the performance boundaries and cost-saving potential of Local SSD-backed node swap, this analysis covers three distinct workload categories: a Traditional Build Workload for CI/CD pipelines, High-Density Browser Sandboxes, and Isolated Python Sandboxes.
| Workload Profile | Baseline Capacity (No Swap) | Local SSD Swap Capacity | Density Improvement |
|---|---|---|---|
| Linux CI/CD Kernel Build | 600 MB RAM Limit | 300 MB RAM Limit | -50% RAM Footprint |
| Headless Chrome (Kata) | 40 Concurrent Pods | 50 Concurrent Pods | +25% Pod Density |
| Headless Chrome (gVisor) | 80 Concurrent Pods | 160 Concurrent Pods | +100% Pod Density |
| Python Sandbox (gVisor) | 80 Concurrent Pods | 240 Concurrent Pods | +200% Pod Density |
1. Traditional workload: Linux kernel build
Before exploring specialized agentic architectures, swap was validated against classic batch workloads by running a complete Linux 6.1.1 kernel build. The kernel compilation process leverages concurrent worker threads, balloons in memory to hold compiled object files, and requires a large memory spike during the brief linking phase.
This workload mirrors the memory behavior of enterprise CI/CD pipelines. Because earlier compiled objects sit inactive in memory while the pipeline progresses, CI/CD jobs frequently hoard unused physical RAM, which makes them well suited to node swap compression.
On a baseline node without swap, the minimum memory limit to prevent an OOM crash during compilation was 600 MB. Routing swap to a Local SSD cut the container memory limit by 50% to 300 MB without incurring any execution slowdown (in fact, it ran cleanly in 374s vs the baseline 433s). However, as an explicit tradeoff, compressing the limit further to 200 MB forced the active working set into swap, causing long I/O wait times and increasing execution time by over 40%. This reinforces that swap serves as an insurance policy for burst memory, not a replacement for active RAM.
2. High-density agent workloads: headless browser runtimes
AI agent workloads frequently require manipulating headless browsers via Chromium. However, trusting external code execution often requires stricter security isolation than standard Linux namespaces. This benchmark cross-evaluated several container runtime environments. The raw logs and testing methodologies for the default runtime are available in the Agent Sandbox GKE Swap directory.
- Unsandboxed baseline limits (
runc): To test the limits of the environment without the overhead of security runtimes, plain runc containers were swept on a c4-standard-32 node (32 vCPU, 120 GB RAM). Without swap, the node exhausted physical memory and failed past 512 pods. Enabling Local SSD swap allowed the node to support 768 concurrent pods. - Advanced security runtimes (for example: gVisor, Kata Containers): Enabling strict security sandboxing increases memory overhead and normally reduces pod density. However, memory swap naturally absorbs this overhead penalty. Without swap, a gVisor environment hit a hard limit at 80 pods. Local SSD swap doubled that capacity, which allowed 160 concurrent gVisor pods on a single node. Similarly, Kata Containers microVMs exhausted physical RAM at 40 concurrent pods without swap, but using GCP Local SSD swap expanded this to 50 stable Kata microVMs before CPU saturation.
At these maximum densities, the per-pod latency increase is driven mainly by pods competing for CPU, not by swap I/O. An operator tuning for a specific latency target would run at a lower density than the peak numbers here and see a proportionally smaller latency cost.
For a comprehensive architectural breakdown and density metrics for gVisor and Kata, see the Agent Sandbox GKE Swap Runtimes directory.
3. Beyond browsers: sandboxed Python runtimes
The advantages of node swap also extend to untrusted, isolated code-execution environments. This sweep deployed simultaneous Python sandbox sessions analyzing 5 million rows of data from the MovieLens 20M dataset, requiring a ≃375 MiB resident memory footprint per execution. The in-depth scaling results and deployment code for this sweep can be reviewed in the Agent Sandbox GKE Swap Python Density directory.
Without swap, heavy concurrent bursts exhausted physical memory, causing the node to hit a hard RAM limit and fail at 80 concurrent sessions. Enabling Local SSD swap offloaded dormant anonymous memory, freeing up physical RAM and preserving the node's page cache. This allowed the node to scale to 240 concurrently isolated Python sandboxes-a 3× density improvement. As with the browser workloads, the latency rise at peak density comes mainly from the sandboxes competing for CPU rather than from swap itself.
Density Benchmarks with and without Node Swap
How to use it
If you manage Kubernetes infrastructure for developer environments, browser testing farms, JVM applications, or AI execution runtimes, leveraging Local SSD swap can multiply your density efficiency.
In Kubernetes v1.34+, node swap is Generally Available. You enable it via the kubelet configuration:
kind: KubeletConfiguration
apiVersion: kubelet.config.k8s.io/v1beta1
failSwapOn: false
memorySwap:
swapBehavior: LimitedSwap
Pairing this upstream configuration with your cloud provider's high-speed local disk gives you dynamic memory balancing. For example, this is natively supported on Google Kubernetes Engine via Node Memory Swap configured on Local SSD profiles.
To get the benefits, configure your workloads with Burstable QoS: set your container's memory limits higher than its requests. The node automatically rations fast swap space based on idle application memory usage while keeping active processes responsive.
Conclusion
As the Kubernetes ecosystem transitions into the agentic era, administrators face a growing conflict between finite hardware memory limits and the bursty behavior of AI workloads. Frameworks like Agent Sandbox provide the security isolation required for running untrusted agents, but that isolation traditionally demands large amounts of idle memory overhead.
By configuring the kubelet with LimitedSwap and routing it to Local SSDs, you can mitigate this conflict. Fast swap offloads the dormant states of idle agents, allowing you to increase pod density and node utilization on the same infrastructure without compromising security boundaries.
05 Oct 2026 6:00pm GMT
22 Sep 2026
Kubernetes Blog
Spotlight on SIG Apps
As Kubernetes adoption has grown, the conversation has shifted beyond running containers to managing increasingly complex application lifecycles. Modern platforms support stateless web services, stateful databases, batch processing, AI workloads, and platform services. At the same time, they must remain reliable during upgrades, scaling events, and infrastructure failures.
Every Kubernetes user relies on SIG Apps, whether they realize it or not. Deployments, StatefulSets, DaemonSets, Jobs, and CronJobs form the foundation of how applications are deployed, updated, scaled, and operated across the Kubernetes ecosystem.
SIG Apps is focused on improving workload resilience, refining application lifecycle management, and addressing the operational challenges that emerge when applications encounter node failures, rollout disruptions, and increasingly complex infrastructure environments.
In this spotlight, we sit down with two of the three SIG Apps chairs Janet Kuo and Maciej Szulik to discuss the evolution of Kubernetes workload management, the challenges of balancing application reliability with operational simplicity, and the future of application lifecycle management within one of Kubernetes' most influential Special Interest Groups.
Introducing SIG Apps
Natalie Fisher: Can you introduce yourself, your role, and how you got involved in SIG Apps?
Janet Kuo: I'm a Senior Staff Software Engineer at Google and have been a Kubernetes maintainer since 2015, joining the community just as we were racing toward the 1.0 launch. In those early days, my focus was on building the core Workloads API, specifically developing controllers like Deployment, ReplicaSet, StatefulSet, and DaemonSet, defining their rollout behaviors, and bringing them from initial designs to GA. That hands-on work was my entry point into SIG Apps.
Since then, I've stayed deeply involved in both the technical and community sides of Kubernetes. I have led SIG Apps as Co-Chair and Tech Lead since 2019. Currently, in addition to maintaining the workloads API, I am driving new subprojects like the Agent Sandbox to ensure Kubernetes is ready for next-generation agentic and AI workloads.
Maciej Szulik: I started contributing to Kubernetes all the way back in 2014. Since then, I've worked across various areas of the project: controllers, kubectl, and apimachinery, which eventually led me to become one of the Chairs and Tech Leads for SIG Apps. My current focus is reliability of the workload controllers under the SIG Apps umbrella and stability and ease of use of kubectl as part of my SIG CLI Tech Lead role. I also care about overall community health and growth as part of my Steering Committee role. Outside of Kubernetes, I work as a Staff Platform Engineer at Defense Unicorns, where I'm helping make Kubernetes more airgap-native with a project called zarf.
The problem and the solution
SIG Apps is responsible for the core workload APIs that power how applications run on Kubernetes. From Deployments and StatefulSets to Jobs and CronJobs, these controllers determine how workloads are created, updated, scaled, and recovered when things go wrong.
As Kubernetes expands to support increasingly diverse workloads - including AI, batch processing, and large-scale distributed applications - SIG Apps continues to evolve these APIs while balancing reliability, backward compatibility, and operational simplicity.
NF: For readers who may not be familiar, what is SIG Apps, and what role does it play within the broader Kubernetes ecosystem?
MS: SIG Apps is the Kubernetes Special Interest Group responsible for the workloads APIs. CronJob and Job help running batch workloads, whereas DaemonSet, Deployment, ReplicaSet, and StatefulSet serve the majority of other applications. More broadly, SIG Apps owns the layer most developers actually touch day-to-day: the controllers that turn a workload specification into running, self-healing pods. It's the group deciding how Deployments roll out, how Jobs retry, how DaemonSets place a pod per node.
JK: Adding to what Maciej described, as the industry shifts, we are seeing a massive demand to run complex, non-traditional workloads like distributed AI training, batch computing, and dynamic agent environments. Our role is expanding: we aren't just maintaining the classic workloads API, but we are actively evolving it and establishing new patterns (like the Agent Sandbox) to make sure Kubernetes remains the best platform for the next generation of workloads, such as AI.
NF: Looking at the workload APIs owned by SIG Apps (Deployments, StatefulSets, DaemonSets, Jobs, and CronJobs), which areas are receiving the most attention from maintainers and contributors?
MS: After a long stretch focused on making batch workloads run smoothly on Kubernetes, we've shifted attention to make sure serving workloads (DaemonSets, StatefulSets, etc) aren't left behind. This means performance and high-scale improvements to rollout and scaling behavior, plus working through our backlog of user-reported issues, prioritizing the ones with the strongest support from the user base.
Current focus areas
As Kubernetes workloads grow in scale and complexity, the challenges facing workload controllers evolve as well. We asked the SIG Apps chairs where contributors are focusing their efforts today and which resilience problems they believe are the highest priorities.
NF: From your perspective, what are the most important workload resilience problems SIG Apps is trying to solve today?
MS: Node lifecycle challenges have come up repeatedly across SIG Apps, SIG Node, and SIG Autoscaling discussions. DaemonSets and Jobs are just where the pain is most visible, since they're the workloads most directly bound to node state. Rather than solve it piecemeal within one SIG, we've settled on spinning up a dedicated Node Lifecycle Working Group to focus on this properly and hopefully land long-term solutions instead of one-off patches.
JK: From an AI perspective, resilience is critical. When you are running a massive distributed LLM training job that spans hundreds of GPUs, a single node failure can halt the entire pipeline. Similarly, if a DaemonSet that runs your logging or GPU monitoring agent gets stuck on a bad node, it impacts the entire cluster's health.
In addition to the work in the Node Lifecycle WG to handle infrastructure-level degradation, SIG Apps is addressing this at the orchestration layer through subprojects like JobSet (for distributed training) and LeaderWorkerSet (LWS) (for sharded LLM inference). These APIs introduce patterns like "all-or-nothing" failure handling, where a single pod or job failure triggers a coordinated group-level restart to resume from the last clean checkpoint, rather than letting stuck workloads hang in an inconsistent state.
Real-world impact
The work happening within SIG Apps extends far beyond controller implementations and API design. We wanted to understand what these improvements mean in practice for platform teams operating Kubernetes clusters in production.
NF: For platform teams operating Kubernetes in production, what practical improvements would they notice if the node lifecycle and workload resilience work currently under discussion is successfully delivered?
MS: I'm mostly looking from the sidelines, the folks actually in the Node Lifecycle Working Group would give you a sharper answer. But from where I sit, I'm hoping their work translates into fewer 3am pages that turn out to be "a DaemonSet rollout got stuck because node X was flaky, and someone had to manually cordon/delete/restart to unstick it."
JK: +1 to what Maciej said, and beyond reducing manual intervention, platform teams will also see much better resource predictability and cost efficiency. For example, in AI workloads where GPU idle time is extremely expensive, having Kubernetes automatically detect a degraded node and reschedule the training coordinator or agent before the job crashes means less wasted compute and more stable job execution.
Challenges and trade-offs
Evolving APIs that millions of workloads rely on requires careful engineering and even more careful decision-making. We asked the SIG Apps chairs about the technical and operational trade-offs they weigh when introducing changes to Kubernetes' core workload controllers.
NF: What are some of the hardest technical or operational trade-offs SIG Apps encounters when evolving core workload controllers?
MS: Honestly, a few tensions keep coming up: how aggressively a controller should give up on stuck pods, and what signals it actually needs to make that call correctly. At the same time, we always have to think about backward compatibility. Deployment, DaemonSet, and Job behavior has been depended on for a decade [by Kubernetes users, tooling, automation, and higher-level controllers], so even a change that's clearly "more correct" can break automation people built around the old behavior without meaning to.
JK: One of our hardest trade-offs is resisting the urge to make "elegant" design changes that break backward compatibility. Instead, we have to design opt-in features that let users adopt new behaviors without forcing them on legacy workloads. When we need to support completely new paradigms, we prefer introducing them as CRDs first rather than bloating the core APIs, like we are doing with Agent Sandbox, JobSet, and LWS.
Looking ahead
While much of SIG Apps' work focuses on maintaining the stability of existing workload APIs, the group is also shaping the future of Kubernetes through new enhancements and proposals. We concluded by asking about one proposal that recently returned to active development and what it represents for the future of workload management.
NF: The SIG recently discussed reviving KEP-4443 with a target release of Kubernetes 1.38. What opportunities or challenges does this proposal aim to address, and why is now the right time to revisit it?
KEP-4443 addresses a small but real gap in the Job API: a PodFailurePolicy can be configured to add a condition reason to the JobFailed condition, but different pod failure policy rules targeting different container exit codes all produce that same generic reason. The proposal is simple: an optional Name field on each PodFailurePolicyRule, which gets appended to the JobFailed condition reason, so higher-level tools like JobSet can finally react differently depending on which rule triggered the failure.
As for timing, the answer is as simple as it always is in open source: we lost the original contributor who was driving this. Now we've got someone new interested in picking it up, that's why we're targeting the next release.
Getting Involved
NF: For someone interested in contributing to SIG Apps, where would you recommend they start, especially if they are not yet a Kubernetes maintainer?
MS: The best place to start is the #sig-apps slack channel and our regular SIG Apps meetings. We've all started there, and if it feels intimidating, or nobody replies right away, that's completely normal. Everyone's busy. It's not personal.
JK: In addition to what Maciej answered, I'd suggest looking at our newer subprojects and initiatives. Contributing to stable APIs like Deployment or StatefulSet can be daunting because the barrier for making changes is very high due to backward compatibility, and there is much less low-hanging fruit.
If you are new to the community, projects like the Agent Sandbox are fantastic entry points. They are actively evolving, have a friendly group of maintainers, and offer plenty of greenfield development opportunities where you can make a significant impact quickly.
Summary
SIG Apps has shaped how Kubernetes applications are deployed and operated since the project's earliest days. While users often interact with Deployments, StatefulSets, Jobs, and DaemonSets without thinking about the controllers behind them, the work within SIG Apps continues to shape the reliability and scalability of workloads across the Kubernetes ecosystem.
From improving workload resilience and node lifecycle behavior to enabling new patterns for AI and distributed computing, the SIG is evolving Kubernetes while remaining committed to one of the project's core principles: preserving the stability and backward compatibility that users depend on. Whether you're interested in core workload APIs, emerging projects like Agent Sandbox, or helping improve the operational experience of Kubernetes users everywhere, SIG Apps offers many opportunities to get involved.
22 Sep 2026 6:00pm GMT
21 Sep 2026
Kubernetes Blog
Kubernetes v1.37: Tracking When a PersistentVolumeClaim Was Last Used (Beta)
Kubernetes v1.37 promotes the PersistentVolumeClaimUnusedSinceTime feature gate to Beta (enabled by default). With this feature, the PersistentVolumeClaim (PVC) protection controller adds an Unused condition to each PVC, telling you whether any running pod currently references it - no custom tooling or cross-referencing required.
For the API definition of PVC conditions, see the PersistentVolumeClaim API reference. Read on to learn how the Unused condition works and how to use it.
Why track PVC usage?
In large-scale Kubernetes clusters, it is common for users to create PVCs and then delete the associated pods without cleaning up the storage, because Kubernetes does not automatically delete PVCs when their pods are removed (to protect against accidental data loss). Over time, these orphaned PVCs may accumulate, silently consuming storage capacity and driving up cloud costs.
Before Kubernetes v1.37, it was easy to identify an unused PersistentVolume, but much harder to determine whether a PVC was still being used. Doing so required cross-referencing pods, PersistentVolumes, and PVCs over a potentially large window of time. Administrators often resorted to custom monitoring pipelines or scripts to answer a seemingly simple question: "Is anything actually using this volume?"
The PersistentVolumeClaimUnusedSinceTime feature solves this by making the answer available natively in the PVC status. Once the feature is enabled, every PVC gets an Unused condition managed by the PVC protection controller.
User stories
- Storage administrator: "I want to know which PVCs in my cluster are not being used by any pod so I can safely identify orphaned volumes and schedule them for deletion."
- DevOps engineer: "I want to list PVCs that have the
Unusedcondition set toTrueso I can automate cleanup in development environments."
How does it work?
The PVC protection controller - which already watches pods to enforce the storage object in use protection - now also manages a new Unused condition on PVCs.
The condition works as follows:
| Scenario | Condition status | Reason |
|---|---|---|
| No non-terminal pods reference the PVC | Unused=True |
NoPodsUsingPVC |
| At least one running or pending pod references the PVC | Unused=False |
PodUsingPVC |
A few details worth noting:
- Terminated pods don't count: A pod that has completed (phase
SucceededorFailed) does not keep the PVC marked as in use. This means batch jobs withrestartPolicy: Neverwon't prevent the PVC from becomingUnused=Trueafter they finish. - Pending pods do count: Even an unschedulable pod (for example, one with an impossible node selector) still counts as using the PVC. The intent to use the volume is enough.
- Multiple pods: If several pods reference the same PVC, the condition transitions to
Unused=Trueonly after the *last" non-terminated pod is removed or terminates.
Using lastTransitionTime to find when a PVC became idle
Like every Kubernetes condition, the Unused condition carries a standard lastTransitionTime field. This means you get a useful bonus for free: when the condition transitions from False to True, the lastTransitionTime records exactly when the PVC became idle. You can use this timestamp to answer questions like "how long has this PVC been sitting unused?" - for example, to find PVCs that have been idle for more than 30 days (see the example query below).
What changed from Alpha to Beta?
Kubernetes v1.36 introduced this feature as Alpha, where you had to enable the PersistentVolumeClaimUnusedSinceTime feature gate explicitly. For Beta in v1.37, the feature gate is enabled by default, and the feature has full end-to-end test coverage.
How to use it
Since the feature is Beta and enabled by default in Kubernetes v1.37, the Unused condition will appear on PVCs automatically. Here is a walkthrough to see it in action:
-
Create a PVC:
apiVersion: v1 kind: PersistentVolumeClaim metadata: name: my-data spec: accessModes: - ReadWriteOnce resources: requests: storage: 1Gi -
After a short time, inspect the PVC conditions:
kubectl get pvc my-data -o jsonpath='{.status.conditions[*]}' | jq .You should see an
Unusedcondition with statusTrueand reasonNoPodsUsingPVC:{ "lastProbeTime": null, "lastTransitionTime": "2026-09-14T12:03:11Z", "message": "No pods are currently referencing this PVC", "reason": "NoPodsUsingPVC", "status": "True", "type": "Unused" } -
Create a pod that uses the PVC:
apiVersion: v1 kind: Pod metadata: name: my-app spec: containers: - name: app image: busybox command: ["sleep", "3600"] volumeMounts: - name: data mountPath: /data volumes: - name: data persistentVolumeClaim: claimName: my-data -
Check the condition again - it should now show
Unused=False:kubectl get pvc my-data -o jsonpath='{.status.conditions[?(@.type=="Unused")].status}'Output:
False -
Delete the pod and wait for the condition to transition back to
Unused=True:kubectl delete pod my-app kubectl get pvc my-data -o jsonpath='{.status.conditions[?(@.type=="Unused")]}'The condition should show
Unused=Truewith reasonNoPodsUsingPVCagain.
Finding unused PVCs across the cluster
To list all PVCs that have been unused for more than 30 days, you can use a command like:
Note:
This command usesjq, a command-line JSON processor.kubectl get pvc -A -o json | jq -r '
.items[]
| select(.status.conditions[]? | select(.type=="Unused" and .status=="True"))
| select(
(.status.conditions[] | select(.type=="Unused") | .lastTransitionTime) as $t
| (now - ($t | fromdateiso8601)) > (30 * 86400)
)
| "\(.metadata.namespace)/\(.metadata.name) unused since \(.status.conditions[] | select(.type=="Unused") | .lastTransitionTime)"
'
What's next?
Depending on feedback and adoption, the Kubernetes project intends to graduate this feature to General Availability (GA) in a future release. If you have feedback on this feature, please open an issue in the kubernetes/kubernetes repository.
To learn more about this enhancement, refer to KEP-5541: PersistentVolumeClaim last used time.
Getting involved
The Kubernetes project always welcomes new contributors. If you would like to get involved, you can join us at SIG Storage.
If you would like to share feedback, you can do so on our public Slack channel (visit https://slack.k8s.io/ for an invitation if you need one).
Special thanks to the contributors who helped design and implement this feature (alphabetical order):
- Arvind Parekh (ArvindParekh)
- Hemant Kumar (gnufied)
- Jan Šafránek (jsafrane)
- Kevin Hannon (kannon92)
- Roman Bednář (RomanBednar)
21 Sep 2026 6:30pm GMT