WebDataset/TFRecord Decoding Crash Mid-Epoch
A large tarball or TFRecord file was partially downloaded, truncated during transfer, or concurrently written to while the training job was reading it. The dataloader hits the truncated EOF mid-stream, throwing a fatal exception that kills the worker.
A large tarball or TFRecord file was partially downloaded, truncated during transfer, or concurrently written to while the training job was reading it.
- Root cause
- A large tarball or TFRecord file was partially downloaded, truncated during transfer, or concurrently written to while the training job was reading it. The dataloader hits the truncated EOF mid-stream, throwing a fatal exception that kills the worker.
- Recommended fix
- Catch and ignore decoding errors in iterable datasets import webdataset as wds dataset = wds.WebDataset(urls).with_handlers(wds.warn_and_continue) Allows the dataloader to skip corrupted records rather than crashing the entire training run.
- How Denpex helps
- Denpex matches WebDataset/TFRecord Decoding Crash Mid-Epoch across every rank in a distributed run and reports which rank failed first, so you act on the initiating node instead of the loudest one.
What this failure is
WebDataset/TFRecord Decoding Crash Mid-Epoch is a Data failure seen during ML training runs. A large tarball or TFRecord file was partially downloaded, truncated during transfer, or concurrently written to while the training job was reading it. The dataloader hits the truncated EOF mid-stream, throwing a fatal exception that kills the worker. Common tags: Dataset Corruption.
Is this what broke your run? Paste your log.
You're reading about WebDataset/TFRecord Decoding Crash Mid-Epoch. Paste your own crash log or traceback below and get the real root cause for YOUR run, not this generic entry. No account, no card. Logs are masked at ingress and never saved to account history.
Want 14 days on the Scale plan?
Request an evaluation code. A verified workplace organization activates up to 50 diagnoses a day, alerts, history, and follow-up questions. No credit card or automatic subscription.
Why it happens (the mechanism)
Because training starts fine, engineers often suspect hardware instability or NaN losses, rather than static data corruption deep in a multi-terabyte dataset.
What you'll observe
- tarfile.ReadError: unexpected end of data
- tf.errors.DataLossError: corrupted record at XXX
Common symptoms and what they mean
| Symptom | Why it happens |
|---|---|
| Training runs fine for several hours, then suddenly crashes. | A large tarball or TFRecord file was partially downloaded, truncated during transfer, or concurrently written to while the training job was reading it. The dataloader hits the truncated EOF mid-stream, throwing a fatal exception that kills the worker. |
| Fails consistently at a specific epoch step. | A large tarball or TFRecord file was partially downloaded, truncated during transfer, or concurrently written to while the training job was reading it. The dataloader hits the truncated EOF mid-stream, throwing a fatal exception that kills the worker. |
Which systems are affected
- WebDataset
- PyTorch
- TFRecord
How to confirm this is the problem
Use this checklist to test the hypothesis against a small reproduction. No single line proves the root cause, so preserve the preceding events and compare one variable at a time.
- ✓Identify the exact file shard being read when the crash occurs.
- ✓Run a checksum/validation pass over all dataset files: `md5sum -c manifest.txt` or `tar -tzf data.tar`
The fix and the prevention pattern
The root cause is on this page and stays free. A free account adds the exact remediation steps, saved history, and the fix on every entry in the encyclopedia.
Sign up free. Unlock the full analysisNo credit card. Daily allowance follows verified trust tier. Instant access.
Diagnose this failure in VS Code
Select the traceback or open the failed terminal, then run Denpex locally to see the initiating rank, collateral failures, exact fix, and verification command without uploading the log.
Install the free VS Code extensionRoot cause
- A large tarball or TFRecord file was partially downloaded, truncated during transfer, or concurrently written to while the training job was reading it. The dataloader hits the truncated EOF mid-stream, throwing a fatal exception that kills the worker.
The fix and how to prevent it
Evaluate Denpex on your own logs
Request a Scale evaluation code. A verified workplace organization activates 14 days with up to 50 diagnoses a day. Every account keeps its current diagnosis allowance and gets a verification path. No card or automatic subscription.
References
Don't just read the fix, diagnose your run
The encyclopedia tells you what went wrong. Denpex tells you what went wrong in YOUR training run. With your logs, your config, and your stack.