Skip to content

WebDataset/TFRecord Decoding Crash Mid-Epoch

A large tarball or TFRecord file was partially downloaded, truncated during transfer, or concurrently written to while the training job was reading it. The dataloader hits the truncated EOF mid-stream, throwing a fatal exception that kills the worker.

Quick answer

A large tarball or TFRecord file was partially downloaded, truncated during transfer, or concurrently written to while the training job was reading it.

Root cause
A large tarball or TFRecord file was partially downloaded, truncated during transfer, or concurrently written to while the training job was reading it. The dataloader hits the truncated EOF mid-stream, throwing a fatal exception that kills the worker.
Recommended fix
Catch and ignore decoding errors in iterable datasets import webdataset as wds dataset = wds.WebDataset(urls).with_handlers(wds.warn_and_continue) Allows the dataloader to skip corrupted records rather than crashing the entire training run.
How Denpex helps
Denpex matches WebDataset/TFRecord Decoding Crash Mid-Epoch across every rank in a distributed run and reports which rank failed first, so you act on the initiating node instead of the loudest one.
Data#Dataset Corruption

What this failure is

WebDataset/TFRecord Decoding Crash Mid-Epoch is a Data failure seen during ML training runs. A large tarball or TFRecord file was partially downloaded, truncated during transfer, or concurrently written to while the training job was reading it. The dataloader hits the truncated EOF mid-stream, throwing a fatal exception that kills the worker. Common tags: Dataset Corruption.

Live diagnosis, no signup

Is this what broke your run? Paste your log.

You're reading about WebDataset/TFRecord Decoding Crash Mid-Epoch. Paste your own crash log or traceback below and get the real root cause for YOUR run, not this generic entry. No account, no card. Logs are masked at ingress and never saved to account history.

training_logs.txt
No log to hand? Try one:

3 free diagnoses/day

Want 14 days on the Scale plan?

Request an evaluation code. A verified workplace organization activates up to 50 diagnoses a day, alerts, history, and follow-up questions. No credit card or automatic subscription.

Evaluate one incident

Why it happens (the mechanism)

Because training starts fine, engineers often suspect hardware instability or NaN losses, rather than static data corruption deep in a multi-terabyte dataset.

What you'll observe

  • tarfile.ReadError: unexpected end of data
  • tf.errors.DataLossError: corrupted record at XXX

Common symptoms and what they mean

SymptomWhy it happens
Training runs fine for several hours, then suddenly crashes.A large tarball or TFRecord file was partially downloaded, truncated during transfer, or concurrently written to while the training job was reading it. The dataloader hits the truncated EOF mid-stream, throwing a fatal exception that kills the worker.
Fails consistently at a specific epoch step.A large tarball or TFRecord file was partially downloaded, truncated during transfer, or concurrently written to while the training job was reading it. The dataloader hits the truncated EOF mid-stream, throwing a fatal exception that kills the worker.

Which systems are affected

  • WebDataset
  • PyTorch
  • TFRecord

How to confirm this is the problem

Use this checklist to test the hypothesis against a small reproduction. No single line proves the root cause, so preserve the preceding events and compare one variable at a time.

  • Identify the exact file shard being read when the crash occurs.
  • Run a checksum/validation pass over all dataset files: `md5sum -c manifest.txt` or `tar -tzf data.tar`

The fix and the prevention pattern

The root cause is on this page and stays free. A free account adds the exact remediation steps, saved history, and the fix on every entry in the encyclopedia.

Sign up free. Unlock the full analysis

No credit card. Daily allowance follows verified trust tier. Instant access.

Diagnose this failure in VS Code

Select the traceback or open the failed terminal, then run Denpex locally to see the initiating rank, collateral failures, exact fix, and verification command without uploading the log.

Install the free VS Code extension

Root cause

  • A large tarball or TFRecord file was partially downloaded, truncated during transfer, or concurrently written to while the training job was reading it. The dataloader hits the truncated EOF mid-stream, throwing a fatal exception that kills the worker.

The fix and how to prevent it

Evaluate Denpex on your own logs

Request a Scale evaluation code. A verified workplace organization activates 14 days with up to 50 diagnoses a day. Every account keeps its current diagnosis allowance and gets a verification path. No card or automatic subscription.

We send a single-use code tied to that address. Static provider and TLD rules do not reject valid addresses. Account trust determines the benefit after signup.

Don't just read the fix, diagnose your run

The encyclopedia tells you what went wrong. Denpex tells you what went wrong in YOUR training run. With your logs, your config, and your stack.