Skip to content

Python 3.13 cannot pickle code objects, so a distributed job reports the wrong failure

An interpreter upgrade changes what pickle will serialise, and traceback objects stop qualifying. Any machinery that ships an exception between ranks then fails while packing it, so the error the cluster reports is the packing failure and the original fault is never printed.

Quick answer

Python 3.13 will not pickle code objects, and an extracted traceback contains them. The error you are reading is the reporting path failing, not the job. Find the real error in the per-rank logs, and format tracebacks to text before sending them.

Symptom
TypeError: cannot pickle code objects
Root cause
An extracted traceback is not a lightweight record. Its frames reference code objects, and code objects are compiled bytecode with no stable serialised representation, so pickling them was always questionable. The newer interpreter stops permitting it, which is a correctness decision rather than a regression.
Recommended fix
Convert the traceback to text at the point of capture and send the string. Formatting produces exactly what a human needs, and a string pickles everywhere.
How Denpex helps
Denpex matches Python 3.13 cannot pickle code objects, so a distributed job reports the wrong failure across every rank in a distributed run and reports which rank failed first, so you act on the initiating node instead of the loudest one.
Environment#python-3.13#pickle#compatibility#version-mismatch#interpreter-upgrade#gather-object

What this failure is

An interpreter-upgrade failure in which traceback objects can no longer be pickled, so any code path that serialises an exception to move it between processes fails and replaces the original diagnosis with its own.

Live diagnosis, no signup

Is this what broke your run? Paste your log.

You're reading about Python 3.13 cannot pickle code objects, so a distributed job reports the wrong failure. Paste your own crash log or traceback below and get the real root cause for YOUR run, not this generic entry. No account, no card. Logs are masked at ingress and never saved to account history.

training_logs.txt
No log to hand? Try one:

3 free diagnoses/day

Why it happens (the mechanism)

Serialising compiled bytecode was never well defined across interpreter versions, and the newer release stops pretending otherwise. Distributed frameworks had come to rely on it because moving an exception whole is the most convenient way to re-raise it on another rank. The reliance is invisible until something fails, at which point the convenience becomes the failure.

What you'll observe

  • The reported error is about pickling and has nothing to do with what actually went wrong
  • The real exception is destroyed by the reporting path, so there is nothing to search for
  • It only appears on the newer interpreter, so it reads as a framework bug rather than an environment change
  • It surfaces on the error path, which by definition is the least exercised code in the system

Common symptoms and what they mean

SymptomWhy it happens
TypeError: cannot pickle code objects raised while an exception was being propagated between ranksAn extracted traceback is not a lightweight record. Its frames reference code objects, and code objects are compiled bytecode with no stable serialised representation, so pickling them was always questionable. The newer interpreter stops permitting it, which is a correctness decision rather than a regression.
The failure appearing inside object gathering or an object-to-tensor conversion rather than in training codeThe failure lands on the reporting path specifically because that is the only place a traceback is deliberately carried around. Normal operation never pickles one, so an upgrade can pass every test and every training step and still break the moment something else fails.
A test asserting on an error message finding the pickling message instead of the one it expectedThe consequence is a lost diagnosis rather than a lost job. Whatever genuinely failed raised first; the attempt to describe it to the other ranks raised second; and only the second one reaches the log.
Identical code succeeding on the previous interpreter minor versionAn extracted traceback is not a lightweight record. Its frames reference code objects, and code objects are compiled bytecode with no stable serialised representation, so pickling them was always questionable. The newer interpreter stops permitting it, which is a correctness decision rather than a regression.

Which systems are affected

  • Distributed error propagation, where a rank's exception is gathered to rank zero to be re-raised
  • Asynchronous checkpoint save, which reports a worker's failure back to the caller
  • Process-pool executors and any launcher that returns exceptions across a process boundary
  • Anything that stores an extracted traceback for later inspection or logging

How to confirm this is the problem

Use this checklist to test the hypothesis against a small reproduction. No single line proves the root cause, so preserve the preceding events and compare one variable at a time.

  • ✓Extract a traceback and attempt to pickle it in a bare interpreter session. Failing on the new version and succeeding on the old one confirms the interpreter, not the framework.
  • ✓Read the per-rank log files rather than the aggregated output; the original exception was recorded there before the propagation attempt.
  • ✓Check whether the failing frame is inside object gathering or a pool executor's result handling, which is where a traceback would be serialised.

Root cause

  • An extracted traceback is not a lightweight record. Its frames reference code objects, and code objects are compiled bytecode with no stable serialised representation, so pickling them was always questionable. The newer interpreter stops permitting it, which is a correctness decision rather than a regression.
  • The failure lands on the reporting path specifically because that is the only place a traceback is deliberately carried around. Normal operation never pickles one, so an upgrade can pass every test and every training step and still break the moment something else fails.
  • The consequence is a lost diagnosis rather than a lost job. Whatever genuinely failed raised first; the attempt to describe it to the other ranks raised second; and only the second one reaches the log.

The fix and how to prevent it

Searchable error signature

search key
TypeError: cannot pickle code objects
    pickle.Pickler(io.BytesIO()).dump(trace)
[<FrameSummary file /tmp/pickly.py, line 2 in <module>>]
AssertionError: 'fail_once policy triggered failure' not found in 'cannot pickle code objects'

Use this text as a lookup key in logs and upstream issue trackers. It is not presented as a captured customer log. Confirm the cause from your own preceding events, versions, configuration and the cited references.

The fix and the prevention pattern

The root cause is on this page and stays free. A free account adds the exact remediation steps, saved history, and the fix on every entry in the encyclopedia.

Sign up free. Unlock the full analysis

No credit card. Daily allowance follows verified trust tier. Instant access.

Why the recommended fix works

Formatting the traceback at capture time preserves everything a person reads from it, the frames, the lines, the final message, in a form that has no version sensitivity at all. Carrying the type name separately keeps whatever programmatic branching depended on the exception class. Nothing of diagnostic value is lost, and the transport stops depending on an interpreter implementation detail.

Best practices by model family

Model / StackRecommendationNotes
You own the propagation codeFormat the traceback to text at captureStrings pickle on every interpreter version.
Framework owns itPin the interpreter minor versionBuys time without editing library internals.
Debugging right nowRead the per-rank logsThe original exception was written before propagation.
Writing new distributed codeNever send frames or code objectsCarry type name, message and formatted text.

With the fix vs without the fix

DimensionWith the fixWithout the fix
What the error describesThe reporting pathRead as the job's actual failure
Where the real cause isIn the per-rank logAssumed lost
What to send between ranksFormatted text and a type nameThe live exception object

Diagnostic note

“The expensive part of this is not the fix, it is the hours spent searching for a pickling bug that does not exist while the actual failure sits unread in a per-rank log. If you take one habit from it, take this: when the reported error is about the machinery of reporting, stop reading the aggregated output and go to the individual ranks. The same reflex pays off for launcher timeouts and for exit-code summaries that name no cause.”

Visual fingerprint

The reporting path eats the diagnosis
  rank 3   real failure raised        -> written to rank-3 log   ✓
             |
             v
  wrap exception + traceback
             |
             v
  gather_object -> pickle -> TypeError: cannot pickle code objects
             |
             v
  aggregated output shows ONLY the pickling error
The genuine failure is recorded on its own rank before anything is serialised. The pickling error replaces it only in the aggregated view, which is why the per-rank logs still hold the answer.

Diagnose this failure in VS Code

Select the traceback or open the failed terminal, then run Denpex locally to see the initiating rank, collateral failures, exact fix, and verification command without uploading the log.

Install the free VS Code extension

Frequently asked questions

Questions engineers and on-call staff commonly ask about this failure.

Did my training job fail because of pickle?
No. Something else failed first, and the attempt to describe that failure to the other ranks is what you are reading. The original error is in the per-rank log.
Why did this appear only after an upgrade?
The newer interpreter stops pickling code objects, which an extracted traceback contains. Nothing in your code changed.
Is downgrading the only option?
It is the quickest if the propagation path belongs to a library. If it is your code, formatting the traceback to text before sending it is a permanent fix.
Will I lose detail by sending text?
No. Formatting yields the frames, lines and final message, everything you read from a traceback anyway.

Don't just read the fix, diagnose your run

The encyclopedia tells you what went wrong. Denpex tells you what went wrong in YOUR training run. With your logs, your config, and your stack.