Python 3.15's Copy-and-Patch JIT: Faster Pure-Python Loops, Same NumPy Ceiling

Python 3.15's Copy-and-Patch JIT: Faster Pure-Python Loops, Same NumPy Ceiling

pr0h0•
pythonjitcpythonperformancenumpy
AI Usage (93%)

Python 3.15's Copy-and-Patch JIT: What It Speeds Up and What It Doesn't

Python 3.15 is out, and the coverage leads with a single item: CPython's copy-and-patch JIT has landed in a mainstream release, tagged experimental. Phoronix's "Python 3.15 Released With Experimental JIT Compiler Running Faster" and InfoWorld's "The best new features in Python 3.15," both published 2026-10-09, set that framing. This post works out what the JIT actually speeds up — which pure-Python loops get faster, why NumPy and other C-extension workloads stay at the same performance ceiling, and how to measure the difference on your own machine.

The question worth asking is narrower than "is Python faster now." Which workloads actually move, and which sit behind a ceiling the JIT cannot reach? My position up front: pure-Python bytecode-bound loops get faster, and anything whose hot path lives inside a C extension — NumPy, Polars, PyTorch, Rust extension modules — does not. The JIT optimizes CPython bytecode. It does not optimize code that is not CPython bytecode.

Scope and Sourcing: What the 3.15 JIT Coverage Actually Establishes

Established by the reports: 3.15 exists, it ships an experimental JIT, and the coverage frames it as running faster. Not established: per-workload numbers, whether the JIT is on by default in the builds people will actually install, the target hardware, and the size of "faster" — 5% and 5x both fit that phrasing.

Phoronix and InfoWorld are both secondary here. For mechanics I'm working from CPython's own JIT design material — PEP 744, the source under Tools/jit and Python/jit.c — plus the original copy-and-patch paper. Anything I did not run myself is flagged as reported, not reproduced.

What Actually Shipped in Python 3.15: An Experimental Tier-2 JIT

Two facts matter more than the version number. First, the JIT is experimental, and that word is load-bearing: flag names, trace heuristics, and the generated code can all shift between point releases. Pin python3.15.0 in production and ride to 3.15.4, and you have pinned a version string, not behavior. Second, this is the first time JIT code has landed in a mainstream release instead of living on a development branch. The practical value there is exposure — far more people will build it, run it, and file bugs than ever touched a dev-branch JIT.

Verify Your Python 3.15 Build Before Benchmarking

Do not trust a distribution's "Python 3.15" package to be a JIT build. Check it first.

python3.15 -VV
python3.15 -c "import sysconfig; print({k: sysconfig.get_config_var(k) for k in ('Py_DEBUG', 'Py_GIL_DISABLED')})"
⚠️

--enable-experimental-jit is the 3.13-era configure flag, and 3.13 also exposed a sys._jit accessor for querying JIT state. Do not assume either name survived into 3.15 — confirm against the 3.15 release notes and ./configure --help on your own checkout. I'm calling this out because flag names are exactly the kind of detail that gets copied forward incorrectly in benchmark posts.

While you're there, confirm you are not benchmarking a Py_DEBUG build. Debug interpreters carry assertions and no optimizations; a number from one tells you nothing about the release binary you ship.

How CPython's Copy-and-Patch Tier-2 JIT Works

Three stages, not a single switch you flip.

Tier 0 and Tier 1: Bytecode, Then Specializing Bytecode

Tier 0 is plain bytecode evaluation. Tier 1 is PEP 659's specializing adaptive interpreter: after warmup, instructions like LOAD_ATTR, BINARY_SUBSCR, and CALL rewrite themselves into type-specialized forms backed by inline caches that remember the last types seen. This is why the JIT isn't competing with the specializing interpreter — it consumes it. It compiles the type feedback those caches collected. Without tier 1, tier 2 would be compiling guesses.

Tier 2: Trace Formation

Once a loop is hot enough, the interpreter starts recording. It executes the loop body, follows the linear path taken across a superblock, and emits a trace. That trace is the compilation unit — not the function, not the module. Cold branches and cold functions stay interpreted, which is the entire point of a tiered design.

The Copy-and-Patch Trick: Precompiled Stencils, Patched at Runtime

At CPython build time, clang compiles a set of small C stencils — roughly one per bytecode operation — into object files. At runtime, the JIT copies the stencil for each operation in the trace into a contiguous buffer and patches the holes with constants, addresses, and the entry point of the next stencil. There is no compiler in the interpreter, only precompiled fragments and something that behaves like a tiny linker. Hence the low trace-compilation latency, and hence no C toolchain at runtime. The technique comes from the copy-and-patch paper by Xu and Kjolstad.

Why the Copy-and-Patch Design Is Cheap, and Why It Is Limited

Cheap: no runtime IR passes, no register allocator, negligible JIT overhead, no external toolchain. Limited: each stencil is compiled in isolation. There is no whole-program analysis, no cross-trace specialization, and no register allocation across trace boundaries. A trace compiled for a monomorphic integer loop deoptimizes the moment an object of a different type shows up at the same site.

Measuring Python 3.15 JIT Speedups on Your Own Machine

A speedup claim without conditions is not a result. Reproduce it.

A Reproducible Pure-Python Benchmark Harness

## jitprobe.py
def work(n):
    acc = 0
    counts = {}
    parts = []
    for i in range(n):
        acc += (i * 3) & 0xFFFF
        k = i % 97
        counts[k] = counts.get(k, 0) + 1
        if (i & 1023) == 0:
            parts.append(str(k))
    return acc, len(counts), len("".join(parts))

That mixes the three shapes the JIT is supposed to help: integer arithmetic in a bytecode-dense loop, monomorphic dict access, and string building. Run it against two interpreters:

python3.15 -m pyperf timeit -s "from jitprobe import work" "work(200000)" --rigorous -o jit315.json
python3.14 -m pyperf timeit -s "from jitprobe import work" "work(200000)" --rigorous -o jit314.json
python3.15 -m pyperf compare_to jit314.json jit315.json

## quick and dirty, for a sanity check only
python3.15 -m timeit -s "from jitprobe import work" -n 100 "work(200000)"

The timeit command prints in this shape:

100 loops, best of 5: 12.4 msec per loop
📝

That output block shows the format, not a measurement I took. I don't have a 3.15 JIT build in front of me while writing this, so every speedup figure in the underlying coverage is reported, not reproduced. The harness above is what you run to replace it with a real number.

What to Record So the Benchmark Is Reproducible

FieldWhy it matters
python -VV + build flagsA non-JIT or debug build invalidates the comparison
CPU model and OSTrace heuristics and codegen differ per target
pyperf run countMedian alone hides the tail
Median and percentilesReport the spread, not one run
Warmup lengthA loop too short to escalate never reaches tier 2
Loop lifetimeShort scripts pay trace compilation and never amortize it

Any number you cite that did not come from this setup should carry the label "reported, not reproduced." Mine included.

Where Pure-Python Loops Get Faster With the JIT

The win surface is narrower than "Python is faster":

  • Bytecode-dense numeric loops — integer and float accumulation, counters, index arithmetic.
  • Repeated attribute and index access where the types stay monomorphic.
  • Tight string and dict manipulation inside a hot loop.
  • Interpreter dispatch overhead disappearing inside the trace.

Where the JIT Win Evaporates

  • Polymorphic call sites. The trace specializes on what the caches saw; a second type triggers deoptimization.
  • Loops whose bodies are dominated by calls into C extensions — the trace is short and the time is not in the trace.
  • Short-lived loops that never hit the hotness threshold.
  • Startup-sensitive scripts that pay trace compilation without living long enough to amortize it.

The NumPy Ceiling: Why C-Extension Workloads Don't Get Faster

Here is the central technical position: the JIT optimizes CPython bytecode, and a NumPy ufunc spends its time in compiled C the JIT never touches.

Why the Python/C Boundary Is the Ceiling

np.add(a, b) on two large arrays involves argument conversion, array boxing, dispatch, and GIL handling. All of that is bytecode, all of it is JIT-able — and all of it is O(1) per call. The O(n) work happens on the other side of the C-ABI boundary. At n = 10^6, shaving per-call overhead changes nothing you can measure. Per-operation cost is unchanged. That is arithmetic, not opinion, and it is why I expect the reported 3.15 speedups to cluster in pure-Python benchmarks rather than numeric ones.

Where the JIT Can Still Help Numerically

  • Small arrays with heavy per-call overhead, where the overhead is the workload.
  • Element-wise Python loops over array data that never vectorized.
  • Orchestration glue across many small C calls: batching logic, index math, conditionals.
  • Generator pipelines feeding C extensions, where the Python side is the cost.

Polars, PyTorch, and Rust Extensions Sit Behind the Same Wall

Polars, PyTorch, and PyO3-based modules all put their hot path on the far side of the same boundary. The JIT can speed up the Python that calls them. It cannot speed up the call.

A Short Checklist for Choosing a JIT Target

Workload shapeWhat to do
Pure-Python hot loop, monomorphicLet the JIT try; measure before rewriting
Mixed Python glue around many small C callsRestructure to batch calls first, JIT second
C-extension-bound, large dataStay in NumPy/Polars; the JIT is irrelevant
Startup-critical CLISkip the JIT build for now; compile cost dominates

Caveats and Open Questions About Python 3.15's JIT

  • Trace memory and code size. Traces are allocated at runtime and retained. What that costs for trace-heavy workloads is untested here.
  • Free-threaded builds. How Py_GIL_DISABLED builds interact with tier-2 traces is not something the reports address; I suspect it is complicated, but that is inference, not a finding.
  • Wheel and ABI implications. Whether a JIT build produces wheels installable on non-JIT builds needs confirmation against the release notes.
  • Default-on status. The coverage says experimental. It does not say whether the 3.15 JIT is meant to become default in a later release.
  • Point releases move. Per-release JIT numbers need re-measuring on each point release, not carried forward.

Conclusion

The JIT narrows the gap between pure Python and native code for bytecode-bound loops, and leaves the C-extension ceiling exactly where it was. That is not a disappointment — it is the design working as specified. The practical takeaway is to work out which of the two you are actually optimizing before you spend a sprint on interpreter upgrades. Profile first. If your hot path is Python bytecode, 3.15 is worth a look. If it is inside NumPy, nothing in this release changes your problem.

Further Reading

Share this post

More posts

Comments