Skip to content

Ephemeral Storage Exhaustion by Model Weights

Machine learning workloads often download massive pre-trained model weights (e.g., from Hugging Face) into the default cache directory (like ~/.cache/huggingface) located on the container's root filesystem. This consumes the node's underlying root partition (ephemeral storage). Without explicit limits, a single pod can fill the disk, triggering a kubelet eviction.

Quick answer

Machine learning workloads often download massive pre-trained model weights (e.

Root cause
Machine learning workloads often download massive pre-trained model weights (e.g., from Hugging Face) into the default cache directory (like ~/.
Recommended fix
HF_HOME=/model-cache
How Denpex helps
Denpex matches Ephemeral Storage Exhaustion by Model Weights across every rank in a distributed run and reports which rank failed first, so you act on the initiating node instead of the loudest one.
Storage#Node Eviction

What this failure is

Ephemeral Storage Exhaustion by Model Weights is a Storage failure seen during ML training runs. Machine learning workloads often download massive pre-trained model weights (e.g., from Hugging Face) into the default cache directory (like ~/.cache/huggingface) located on the container's root filesystem. This consumes the node's underlying root partition (ephemeral storage). Without explicit limits, a single pod can fill the disk, triggering a kubelet eviction. Common tags: Node Eviction.

Live diagnosis, no signup

Is this what broke your run? Paste your log.

You're reading about Ephemeral Storage Exhaustion by Model Weights. Paste your own crash log or traceback below and get the real root cause for YOUR run, not this generic entry. No account, no card. Logs are masked at ingress and never saved to account history.

training_logs.txt
No log to hand? Try one:

3 free diagnoses/day

Want 14 days on the Scale plan?

Request an evaluation code. A verified workplace organization activates up to 50 diagnoses a day, alerts, history, and follow-up questions. No credit card or automatic subscription.

Evaluate one incident

Why it happens (the mechanism)

Engineers often monitor GPU VRAM and system RAM, ignoring disk space. The failure happens dynamically during runtime (downloading weights) rather than at scheduling time, making it look like a crash loop rather than a resource capacity issue.

What you'll observe

  • The node was low on resource: ephemeral-storage.
  • Container <name> was using <size>, which exceeds its request of 0.

Common symptoms and what they mean

SymptomWhy it happens
Pod is suddenly evicted with status 'Evicted' during startup or model initialization.Machine learning workloads often download massive pre-trained model weights (e.g., from Hugging Face) into the default cache directory (like ~/.cache/huggingface) located on the container's root filesystem. This consumes the node's underlying root partition (ephemeral storage). Without explicit limits, a single pod can fill the disk, triggering a kubelet eviction.
The host node transitions into a 'DiskPressure' state.Machine learning workloads often download massive pre-trained model weights (e.g., from Hugging Face) into the default cache directory (like ~/.cache/huggingface) located on the container's root filesystem. This consumes the node's underlying root partition (ephemeral storage). Without explicit limits, a single pod can fill the disk, triggering a kubelet eviction.
Other random pods on the same node might also be evicted.Machine learning workloads often download massive pre-trained model weights (e.g., from Hugging Face) into the default cache directory (like ~/.cache/huggingface) located on the container's root filesystem. This consumes the node's underlying root partition (ephemeral storage). Without explicit limits, a single pod can fill the disk, triggering a kubelet eviction.

Which systems are affected

  • Kubernetes Kubelet
  • Container Runtime
  • HuggingFace Transformers

How to confirm this is the problem

Use this checklist to test the hypothesis against a small reproduction. No single line proves the root cause, so preserve the preceding events and compare one variable at a time.

  • kubectl get events --sort-by='.metadata.creationTimestamp' | grep Evicted
  • kubectl describe node <node-name> | grep DiskPressure
  • Check pod storage usage using 'kubectl top pod --containers'

The fix and the prevention pattern

The root cause is on this page and stays free. A free account adds the exact remediation steps, saved history, and the fix on every entry in the encyclopedia.

Sign up free. Unlock the full analysis

No credit card. Daily allowance follows verified trust tier. Instant access.

Diagnose this failure in VS Code

Select the traceback or open the failed terminal, then run Denpex locally to see the initiating rank, collateral failures, exact fix, and verification command without uploading the log.

Install the free VS Code extension

Root cause

  • Machine learning workloads often download massive pre-trained model weights (e.g., from Hugging Face) into the default cache directory (like ~/.cache/huggingface) located on the container's root filesystem. This consumes the node's underlying root partition (ephemeral storage). Without explicit limits, a single pod can fill the disk, triggering a kubelet eviction.

The fix and how to prevent it

Evaluate Denpex on your own logs

Request a Scale evaluation code. A verified workplace organization activates 14 days with up to 50 diagnoses a day. Every account keeps its current diagnosis allowance and gets a verification path. No card or automatic subscription.

We send a single-use code tied to that address. Static provider and TLD rules do not reject valid addresses. Account trust determines the benefit after signup.

Don't just read the fix, diagnose your run

The encyclopedia tells you what went wrong. Denpex tells you what went wrong in YOUR training run. With your logs, your config, and your stack.