Ephemeral Storage Exhaustion by Model Weights
Machine learning workloads often download massive pre-trained model weights (e.g., from Hugging Face) into the default cache directory (like ~/.cache/huggingface) located on the container's root filesystem. This consumes the node's underlying root partition (ephemeral storage). Without explicit limits, a single pod can fill the disk, triggering a kubelet eviction.
Machine learning workloads often download massive pre-trained model weights (e.
- Root cause
- Machine learning workloads often download massive pre-trained model weights (e.g., from Hugging Face) into the default cache directory (like ~/.
- Recommended fix
HF_HOME=/model-cache- How Denpex helps
- Denpex matches Ephemeral Storage Exhaustion by Model Weights across every rank in a distributed run and reports which rank failed first, so you act on the initiating node instead of the loudest one.
What this failure is
Ephemeral Storage Exhaustion by Model Weights is a Storage failure seen during ML training runs. Machine learning workloads often download massive pre-trained model weights (e.g., from Hugging Face) into the default cache directory (like ~/.cache/huggingface) located on the container's root filesystem. This consumes the node's underlying root partition (ephemeral storage). Without explicit limits, a single pod can fill the disk, triggering a kubelet eviction. Common tags: Node Eviction.
Is this what broke your run? Paste your log.
You're reading about Ephemeral Storage Exhaustion by Model Weights. Paste your own crash log or traceback below and get the real root cause for YOUR run, not this generic entry. No account, no card. Logs are masked at ingress and never saved to account history.
Want 14 days on the Scale plan?
Request an evaluation code. A verified workplace organization activates up to 50 diagnoses a day, alerts, history, and follow-up questions. No credit card or automatic subscription.
Why it happens (the mechanism)
Engineers often monitor GPU VRAM and system RAM, ignoring disk space. The failure happens dynamically during runtime (downloading weights) rather than at scheduling time, making it look like a crash loop rather than a resource capacity issue.
What you'll observe
- The node was low on resource: ephemeral-storage.
- Container <name> was using <size>, which exceeds its request of 0.
Common symptoms and what they mean
| Symptom | Why it happens |
|---|---|
| Pod is suddenly evicted with status 'Evicted' during startup or model initialization. | Machine learning workloads often download massive pre-trained model weights (e.g., from Hugging Face) into the default cache directory (like ~/.cache/huggingface) located on the container's root filesystem. This consumes the node's underlying root partition (ephemeral storage). Without explicit limits, a single pod can fill the disk, triggering a kubelet eviction. |
| The host node transitions into a 'DiskPressure' state. | Machine learning workloads often download massive pre-trained model weights (e.g., from Hugging Face) into the default cache directory (like ~/.cache/huggingface) located on the container's root filesystem. This consumes the node's underlying root partition (ephemeral storage). Without explicit limits, a single pod can fill the disk, triggering a kubelet eviction. |
| Other random pods on the same node might also be evicted. | Machine learning workloads often download massive pre-trained model weights (e.g., from Hugging Face) into the default cache directory (like ~/.cache/huggingface) located on the container's root filesystem. This consumes the node's underlying root partition (ephemeral storage). Without explicit limits, a single pod can fill the disk, triggering a kubelet eviction. |
Which systems are affected
- Kubernetes Kubelet
- Container Runtime
- HuggingFace Transformers
How to confirm this is the problem
Use this checklist to test the hypothesis against a small reproduction. No single line proves the root cause, so preserve the preceding events and compare one variable at a time.
- ✓kubectl get events --sort-by='.metadata.creationTimestamp' | grep Evicted
- ✓kubectl describe node <node-name> | grep DiskPressure
- ✓Check pod storage usage using 'kubectl top pod --containers'
The fix and the prevention pattern
The root cause is on this page and stays free. A free account adds the exact remediation steps, saved history, and the fix on every entry in the encyclopedia.
Sign up free. Unlock the full analysisNo credit card. Daily allowance follows verified trust tier. Instant access.
Diagnose this failure in VS Code
Select the traceback or open the failed terminal, then run Denpex locally to see the initiating rank, collateral failures, exact fix, and verification command without uploading the log.
Install the free VS Code extensionRoot cause
- Machine learning workloads often download massive pre-trained model weights (e.g., from Hugging Face) into the default cache directory (like ~/.cache/huggingface) located on the container's root filesystem. This consumes the node's underlying root partition (ephemeral storage). Without explicit limits, a single pod can fill the disk, triggering a kubelet eviction.
The fix and how to prevent it
Evaluate Denpex on your own logs
Request a Scale evaluation code. A verified workplace organization activates 14 days with up to 50 diagnoses a day. Every account keeps its current diagnosis allowance and gets a verification path. No card or automatic subscription.
References
Don't just read the fix, diagnose your run
The encyclopedia tells you what went wrong. Denpex tells you what went wrong in YOUR training run. With your logs, your config, and your stack.