← Runbooks

How to Fix OOMKilled in Kubernetes (2026)

OOMKilled means a container was terminated for using more memory than its limit. How to confirm exit code 137, and how to tell it apart from a Pending pod.

Key Takeaways

  • OOMKilled means the container was terminated because it ran out of memory. The Kubernetes memory docs show the terminated state as reason: OOMKilled with exitCode: 137.
  • A container may use more memory than it requests, but not more than its limit. "A Container is guaranteed to have as much memory as it requests, but is not allowed to use more memory than its limit."
  • Crossing the limit makes the container a candidate for termination. "If a Container allocates more memory than its limit, the Container becomes a candidate for termination. If the Container continues to consume memory beyond its limit, the Container is terminated."
  • The kubelet restarts it if the container can be restarted. "If a terminated Container can be restarted, the kubelet restarts it, as with any other type of runtime failure." That restart loop is why OOMKilled often shows up as CrashLoopBackOff.
  • Pending is a different failure. A memory request larger than any node leaves the pod Pending with FailedScheduling and Insufficient memory. The container never started, so it was not OOMKilled.
  • A container with no memory limit is easier for the node to kill. "A container with no resource limits will have a greater chance of being killed."

OOMKilled is the reason Kubernetes records when it terminates a container for using more memory than its limit. The pod may then sit in CrashLoopBackOff, because the kubelet restarts a restartable container the same way it restarts any other runtime failure. The fix starts by confirming reason: OOMKilled and exitCode: 137, then deciding whether the limit is too low or the process is allocating without bound.

What does OOMKilled mean in Kubernetes?

It means the container was killed for memory, not that the scheduler refused to place the pod. The assign memory resources page states the rule directly: "A Container is guaranteed to have as much memory as it requests, but is not allowed to use more memory than its limit."

The same page describes what happens past the limit: "If a Container allocates more memory than its limit, the Container becomes a candidate for termination. If the Container continues to consume memory beyond its limit, the Container is terminated. If a terminated Container can be restarted, the kubelet restarts it, as with any other type of runtime failure."

The documented signature in kubectl get pod --output=yaml is under lastState.terminated:

reason: OOMKilled
exitCode: 137

kubectl describe nodes records the kill on the node as well. The docs show that event as "Warning OOMKilling Memory cgroup out of memory." kubectl describe pod shows the restart cycle as "Warning BackOff Back-off restarting failed container."

How do I confirm a container was OOMKilled?

Read the terminated state, then the node event. Do not treat a high restart count by itself as proof.

kubectl get pod <pod-name> --output=yaml

Look for reason: OOMKilled and exitCode: 137 on the container that already stopped. Then:

kubectl describe pod <pod-name>
kubectl describe nodes

describe pod shows the backoff warning. describe nodes shows the OOMKilling warning. If those fields are absent, the container exited for a different reason. The previous container log is still worth reading, but an OOM kill often leaves little application output, because the process is terminated rather than exiting on its own.

How is OOMKilled different from Pending?

A pod that cannot be scheduled never reaches OOMKilled. Scheduling uses requests, not limits. The docs are explicit: "Pod scheduling is based on requests. A Pod is scheduled to run on a Node only if the Node has enough available memory to satisfy the Pod's memory request." The pod request is "the sum of the memory requests for all the Containers in the Pod."

When that sum is larger than every node, kubectl get pod shows Pending, and kubectl describe pod shows FailedScheduling with Insufficient memory. Raise the request only if a node can actually hold it, or add capacity. That is not a memory-limit change.

StatusWhat happenedWhere it shows up
OOMKilledThe container ran, then exceeded its memory limitreason: OOMKilled, exitCode: 137, node event OOMKilling
CrashLoopBackOffThe kubelet is waiting between restarts of a container that keeps dyingRestart count climbing, Back-off restarting failed container
PendingThe pod was never scheduledFailedScheduling, Insufficient memory

OOMKilled and CrashLoopBackOff often appear together. The kill is the cause. CrashLoopBackOff is the backoff wrapped around the restart. Steps for the backoff itself are in how to fix CrashLoopBackOff.

What are the common causes of OOMKilled?

The status has one mechanism and a few ways to arrive there.

CauseWhere it shows upTypical fix
Working set above the memory limitOOMKilled while the process is otherwise healthyRaise resources.limits.memory to cover the real working set
Unbounded allocation or a leakMemory climbs until the same limit kills the container againFix the allocation. A higher limit only delays the next kill
No memory limit setThe container can use node memory, and the docs say it has "a greater chance of being killed"Set a limit, or expect the node OOM killer to choose it
Memory request larger than any nodePod stays Pending, not OOMKilledLower the request or add a node that can satisfy it

Limits and requests are set per container with resources.limits.memory and resources.requests.memory. If you omit a limit, one of two things applies, per the same page: the container has no upper bound and can consume node memory, or a namespace LimitRange assigns a default limit.

How do I fix OOMKilled step by step?

Confirm the kill, separate it from Pending, then change the limit or the process. Do not raise the limit until reason: OOMKilled is actually there.

Diagram of an OOMKilled diagnosis. The symptom is a container terminated with reason OOMKilled and exit code 137, because a container is not allowed to use more memory than its limit and the kubelet restarts it if it can. Two checks follow: kubectl get pod -o yaml shows lastState.terminated.reason OOMKilled and exitCode 137, and kubectl describe nodes shows Warning OOMKilling, Memory cgroup out of memory. The signature then splits: OOMKilled means the container ran and exceeded its memory limit, while Pending with FailedScheduling Insufficient memory means the pod never started because the request is larger than any node. Recovery is a restart count that stops climbing while the container stays under its memory limit.

After the change, watch the restart count. A container that stays under its limit stops producing OOMKilled. A container that climbs to the new limit and dies again is leaking or allocating without a bound. The limit was not the bug.

How does an AI SRE help with OOMKilled?

The commands are short. The slow part is telling a limit that is simply too low apart from a leak, and telling both apart from a pod that is Pending because the request cannot be scheduled. Aurora is an open-source AI SRE that reads pod status and node events through Kubernetes tools, then correlates them with recent changes rather than acting blind.

Aurora runs a single investigation agent by default, with an opt-in multi-agent orchestrator, and its powerful tools sit behind an explicit allowlist. It supports sandboxed Kubernetes execution when ENABLE_POD_ISOLATION is set. Any fix is proposed as a pull request a human merges, never an automatic change to a running cluster. For the safety model, see AI agent kubectl safety. For the broader workflow, see AI-powered incident investigation.

The summary

OOMKilled means the container exceeded its memory limit and was terminated with exit code 137. Confirm it in the pod YAML and in the node's OOMKilling event. A Pending pod with Insufficient memory never started, so changing its limit will not schedule it. If the container is restarting, CrashLoopBackOff is the backoff around this kill, not a second root cause.

Try Aurora

To self-host and point Aurora at your own cluster:

git clone https://github.com/Arvo-AI/aurora

Sourcing note. The memory guarantee, the termination rule for a container that exceeds its limit, the kubelet restart sentence, reason: OOMKilled with exitCode: 137, the OOMKilling and BackOff events, the Pending FailedScheduling / Insufficient memory case, and the note on containers with no limit are quoted from the Kubernetes page "Assign Memory Resources to Containers and Pods," verified 28 September 2026. Aurora's defaults are described from its open-source repository.

Frequently Asked Questions