Changing CPU or memory for a running workload has traditionally meant replacing the Pod. That is easy to automate, but replacement still has a cost: connections drain, caches warm again, readiness gates run, and stateful or batch workloads may lose useful progress.
Kubernetes 1.35 graduated in-place Pod resource resize to General Availability. CPU and memory requests or limits can be changed for existing Pods without recreating the Pod or its containers when the resize can be applied in place. The feature is useful precisely because it makes a common operational change less disruptive. It is not an excuse to stop capacity planning.
Know whether the service is CPU-bound, memory-constrained, throttled, or simply slow somewhere downstream.
Change one resource dimension at a time when possible and observe the actual application response.
A technically successful resize can still worsen density, scheduling headroom, latency, or cost.
What changed in Kubernetes 1.35?
The important shift is operational: a resource adjustment no longer necessarily implies a Pod replacement. For workloads where restart cost matters, that gives platform teams another control surface between “leave it alone” and “recreate the workload.”
As of October 2026, the Kubernetes project lists the 1.35 release line as actively supported, with 1.35.9 released on September 15. The project currently schedules 1.35 to enter maintenance mode on December 28, 2026 and reach end of life on February 28, 2027. Version support still matters: a feature being GA does not make an old cluster magically current.
Start with the symptom, not the YAML
Before raising a request or limit, establish what pressure looks like. CPU throttling, sustained CPU saturation, out-of-memory kills, working-set growth, garbage-collection pressure, queue depth, request latency, and downstream wait time tell different stories.
- CPU throttling with spare node capacity: a CPU limit may be constraining useful work.
- High CPU with throughput increasing proportionally: more CPU may buy capacity.
- Memory rising without stabilizing: more memory may only postpone an application leak.
- Low CPU while latency climbs: investigate I/O, locks, databases, queues, and external dependencies before buying compute.
Requests and limits still mean different things
A request participates in scheduling and represents the resource the workload asks Kubernetes to reserve. A limit constrains how much resource a container may consume. Treating both as one “size” hides the reason for the change.
If I am fixing chronic CPU throttling, I want evidence about the limit. If I am improving scheduling guarantees, I care about the request. If I change both together, I record why, because future operators otherwise inherit a number with no decision attached to it.
# Inspect the current resources first
kubectl get pod api-7d8f9c -o jsonpath='{.spec.containers[*].resources}'
# Then inspect runtime behavior and Pod events
kubectl top pod api-7d8f9c
kubectl describe pod api-7d8f9c
Use a resize as an experiment
A good resize has a hypothesis and a stop condition. “Raise CPU from 500m to 750m because p95 latency degrades when throttling exceeds X” is testable. “The service looks busy” is how resource requests slowly become folklore.
- Capture a baseline for latency, throughput, throttling, memory working set, restarts, and node headroom.
- Apply the smallest meaningful change to a limited set of Pods or a non-critical environment first.
- Observe long enough to cover representative traffic, not merely the five quiet minutes after the command succeeds.
- Compare service-level outcomes and cluster-level cost or density.
- Keep the change only if the expected bottleneck moved in the right direction.
Watch the cluster, not just the Pod
Vertical changes affect bin packing. A service can become healthier while the cluster becomes harder to schedule efficiently. Raising requests across many replicas reduces allocatable headroom; raising memory limits can increase the blast radius of simultaneous growth.
That means the rollout dashboard should include pending Pods, node utilization, scheduling failures, eviction signals, and autoscaler behavior alongside application metrics. Production engineering is inconveniently holistic like that.
Do not confuse in-place resize with autoscaling
Horizontal Pod Autoscaling changes replica count. Vertical resource tuning changes resources per Pod. In-place resize makes one vertical operation less disruptive, but the control-loop question remains: which dimension should respond to load?
Stateless request-serving workloads often benefit from horizontal scaling when work parallelizes cleanly. Stateful workers, expensive warm-up paths, or workloads with per-process memory needs may have stronger reasons to adjust vertically. Many real systems use both, with explicit boundaries so two autoscalers do not fight over the same signal.
There is another new restart primitive, but it is different
Kubernetes 1.35 also introduced an alpha RestartAllContainers action behind the RestartAllContainersOnContainerExits feature gate. It can restart all containers in a Pod in place while preserving the Pod identity, IP, sandbox, attached devices, and volumes. That is a recovery primitive, not a resizing mechanism, and it has different lifecycle implications, including re-running init containers.
I would keep that alpha feature out of a routine production resizing plan unless the workload has a concrete recovery use case and the team understands the feature-gate and lifecycle tradeoffs.
A production checklist
- Confirm the cluster version and provider support before assuming the capability is available.
- Record the bottleneck and the expected effect of the resize.
- Separate request changes from limit changes conceptually, even if both are edited.
- Measure application latency, throughput, CPU throttling, memory behavior, and restart counts.
- Measure node headroom, pending Pods, evictions, and scheduling pressure.
- Roll out gradually and keep the previous resource values available for rollback.
- Update the workload specification or delivery source of truth so a later deployment does not silently restore stale values.
- Revisit autoscaling thresholds after a material per-Pod capacity change.