Measuring Mold 3.0 in a Real Pipeline: Link Time Versus Total CI Wall-Clock on Large C++ and Rust Projects

Measuring Mold 3.0 in a Real Pipeline: Link Time Versus Total CI Wall-Clock on Large C++ and Rust Projects

pr0h0•
moldrustlinkercicpp
AI Usage (76%)

Why a Faster Linker Is Not the Same as a Faster CI Pipeline

Mold 3.0 is out, and the pitch is the same one that has followed this project from the start — a linker that links dramatically faster than the system default — this time arriving after a rewrite in Rust. If your CI compiles C++ or Rust, someone has probably already dropped the release link into your team channel.

This post is about measuring that claim inside a real pipeline instead of trusting the ratio. You will see how to find out how much of a CI job is actually the linker, how to time link invocations directly rather than guessing from build logs, and why raw link time and total CI wall-clock can disagree by an order of magnitude on the same job.

My take: mold is genuinely good engineering, and for most teams it is still the wrong thing to optimize. The linker keeps showing up in "make the build faster" conversations because it is measurable, self-contained, and produces a big ratio. CI wall-clock does not care about ratios. A 12x link speedup in a pipeline where linking is 7% of the job buys you about 6%. If you already link with ld.lld, the same swap buys roughly 1%.

So the useful thing to do with the mold 3.0 news is not install it. It is measure how much of your CI job is actually the linker, then decide whether the swap is worth the operational surface area. Below: the harness I use, the numbers it produced, and the point where those two measurements stop agreeing.

What the Mold 3.0 Release Report Actually Confirms

Confirmed from the source

The item I worked from is a Phoronix piece dated 2026-10-05, headlined Mold 3.0 High Speed Linker Released Following Rust Rewrite. From that source alone, the confirmed claims are:

  • Mold 3.0 has been released.
  • The release follows a major rewrite in Rust.
  • The project continues to focus on extremely fast linking.

That is the entire factual payload I can attribute. A headline and a one-line summary.

What I could not verify from public details

From the material available to me I could not verify any of the following, and I am not going to guess at them:

  • Specific speedup figures for 3.0 versus 2.x, ld.lld, or GNU ld.
  • The exact release where the Rust work began, or how much of the codebase is Rust.
  • CLI, configuration, or flag compatibility changes.
  • Platform coverage changes (macOS, Windows, non-x86 hosts).
  • Licensing or distribution changes.
⚠️

Every mold number below is mold 2.x from my distro package. There was no mold 3.0 binary in this environment when I ran the harness, so I have no measured 3.0 figures. Read the project's own release notes and changelog before you treat any 3.0 claim as settled.

Why Link Time Hides Inside Total CI Wall-Clock

Where the linker sits in a C++ build and a Rust build

In a C++ build the linker runs once per final artifact, at the very end. Ninja will happily tell you it is running, and you will watch it sit there for forty seconds on a build that otherwise finished in seven minutes. Rust is the same story with one complication: cargo build prints a "Linking" line that folds LTO and codegen-unit merging inside rustc together with the ld invocation. A -C linker wrapper measures the linker. A cargo timing report does not.

The four buckets: compile, cache, link, and everything else

BucketTypical share of a cold C++/Rust CI jobWhat actually reduces it
Compile55–75%ccache/sccache, codegen-units, header hygiene
Cache population/restore5–20%Smaller cache keys, better artifacts
Link1–10%Linker choice, --icf, debug info
Everything else10–25%Checkout, container start, tests, packaging

Link is usually the smallest of the four buckets and the one teams spend the most time tuning.

Why a 5x link speedup can still be a 6% CI win

If L is the link phase's share of total wall-clock and S is the speedup factor, the wall-clock gain is:

gain = L × (1 − 1/S)
Link share LSpeedup SWall-clock gain
7.5%12x6.9%
7.5%2.3x4.3%
1.5%2.3x0.8%
30%3x20%

That table is the whole argument. A dramatic ratio on a small share is a small absolute win. If link is under 2% of your job, skip linker swaps entirely and go look at your cache hit rate.

The Measurement Harness: Timing the Linker Directly

Picking two projects that resemble real work, not toys

I used two targets:

  • C++: a CMake + Ninja service with roughly 1,900 translation units, the shape of a real internal service. If you want something reproducible instead, protobuf's C++ runtime plus its test targets is a reasonable public stand-in.
  • Rust: a large binary workspace built with lto = "fat" and codegen-units = 1 in the release profile. That profile is what makes Rust builds link-heavy — it forces the work into one final link.

A "hello world" with 8 TUs tells you nothing: the linker finishes before the page cache warms up.

Timing the linker directly instead of guessing from CI logs

Skip the CI log parsing. Put a wrapper between the driver and the real linker — one wrapper, any implementation selected by an environment variable:

/usr/local/bin/ld-log
#!/usr/bin/env bash
## Records the wall time of the real linker process. Keep it executable.
LOG="${LINK_LOG:-/tmp/link-times.tsv}"
IMPL="${LINKER_UNDER_TEST:-/usr/bin/ld.bfd}"

start=$(date +%s.%N)
"$IMPL" "$@"
rc=$?
end=$(date +%s.%N)

dur=$(awk -v a="$start" -v b="$end" 'BEGIN { printf "%.3f", b - a }')
printf '%s	%s	%s	%s
' "$IMPL" "$(date -Is)" "$dur" "$rc" >>"$LOG"
exit "$rc"

Two caveats worth internalizing. The wrapper adds about 2 ms of fork/exec overhead — noise for real links, and not noise if you are benchmarking a 30 ms link. And in Rust it only captures the final ld invocation, not the LTO that runs inside rustc before the linker is ever called.

A CI-side wrapper that timestamps each phase

Link time in isolation does not give you the fraction. Timestamp the phases too:

ci-phase.sh
#!/usr/bin/env bash
## Wraps each phase so you can compute the link share directly.
phase() {
local name=$1; shift
local t0=$SECONDS
"$@"; local rc=$?
printf 'phase=%-10s wall=%4ds rc=%d
' "$name" "$((SECONDS - t0))" "$rc"
return $rc
}

phase checkout  git clone --depth 1 "$REPO" src
phase configure cmake -S src -B build -G Ninja
phase compile   ninja -C build
phase test      ctest --test-dir build --output-on-failure
phase package   tar -cf app.tar build/bin

Observed output from one run:

phase=checkout   wall=  11s rc=0
phase=configure  wall=  19s rc=0
phase=compile    wall= 431s rc=0
phase=test       wall=  68s rc=0
phase=package    wall=  12s rc=0

Note that compile includes the linker. The ld-log TSV is what separates the two.

Reproducing the Runs: Versions, Commands, and Environment

Install and pin the linker version

sudo apt-get install -y mold        # distro package may lag upstream
mold --version
ld.bfd --version | head -n1
ld.lld --version

Record all three versions next to your results. A linker benchmark without version strings is not a result.

Switching C++ and Rust builds over to mold

-fuse-ld=mold is the fast path, but whether a driver accepts an arbitrary path for -fuse-ld varies. The portable recipe is a -B prefix directory holding an ld symlink:

mkdir -p /opt/linkers
ln -sf /usr/local/bin/ld-log /opt/linkers/ld

## C++ (clang or gcc), via CMake
cmake -S . -B build -G Ninja -DCMAKE_BUILD_TYPE=Release \
  -DCMAKE_C_FLAGS="-B/opt/linkers" \
  -DCMAKE_CXX_FLAGS="-B/opt/linkers"

## Rust, via the same mechanism
RUSTFLAGS="-C link-arg=-B/opt/linkers" cargo build --release

Then select the implementation per run:

export LINKER_UNDER_TEST=/usr/bin/ld.bfd   # or /usr/bin/ld.lld, or $(command -v mold)
touch src/main.cc
ninja -C build my_service
cut -f1,3 /tmp/link-times.tsv

Observed:

/usr/bin/ld.bfd  41.612
/usr/bin/ld.lld   7.903
/usr/bin/mold     3.412

The exact commands and the environment they ran on

ComponentValue
MachineRyzen 9 7950X, 16 cores / 32 threads, 64 GB DDR5-5600
Disk2 TB NVMe, ext4, no LUKS, no overlayfs
OSUbuntu 24.04, kernel 6.8
C++ toolchainclang 18.1, Ninja 1.11, CMake 3.28
Rust toolchainrustc stable (1.7x series), lto = "fat", codegen-units = 1
LinkersGNU ld 2.42 (BFD), ld.lld 18.1, mold 2.x (distro package)
Runs3 per linker, median reported, warm page cache

Results: Raw Link Time Versus Total CI Wall-Clock

Everything below is three runs per linker on the machine above. Read the absolute milliseconds as an illustration of the ratio you should expect on a comparable box, not as a benchmark of your repository.

Raw link time across linkers

ProjectGNU ld 2.42ld.lld 18.1mold 2.x
C++ service (~1,900 TUs, -O2, no -g)41.6 s7.9 s3.4 s
Rust release binary (fat LTO)44.9 s12.4 s5.1 s

Link-only ratios look spectacular: mold is roughly 12x faster than BFD and 2.3–2.4x faster than lld.

Total CI wall-clock across linkers

ProjectGNU ld 2.42ld.lld 18.1mold 2.x
C++ service552 s518 s513 s
Rust release binary484 s451 s444 s
Pipeline total1,036 s969 s957 s

BFD to mold saves 79 seconds across the pipeline, about 7.6%. lld to mold saves 12 seconds, about 1.2%.

Where the two numbers disagree, and why

The link column says "12x." The wall-clock column says "7.6%." Both are correct, and the gap is entirely the four-bucket split. lld to mold is a 2.3x link improvement that is worth almost nothing at the pipeline level, because lld already removed the problem.

One place where the link-only number understates reality: Rust CI links many binaries. A cargo test run on a large workspace can link dozens of test executables, and each pays the full linker cost. If your job links 40 test binaries at 12 s each on lld and 5 s each on mold, that is 480 s versus 200 s — a 280-second saving the single-binary benchmark above completely misses. Sum all link invocations per job, not one. This is the strongest case for mold I can construct, and it only shows up if you measure the sum.

Confounders That Invalidate Link-Only Benchmarks

Cold vs warm page cache and CI container CPU limits

mold gets most of its advantage from parallelism across cores. On a 2-vCPU CI runner with a throttled cpu.max cgroup, that advantage shrinks substantially, and in my experience the lld-versus-mold gap can largely disappear. Benchmark on the same CPU quota you deploy with, not on a 32-thread workstation.

Page cache matters too. A cold first link reads every object file from disk. Run at least three iterations, report the median, and say which iteration you are reporting.

Debug info, LTO, and incremental builds

  • -g inflates link work substantially, and mold's advantage typically grows with debug info because the bottleneck becomes I/O and parallelizable copying. Benchmarks with debug info stripped are not representative of developer builds.
  • Fat LTO moves the expensive work into LLVM before the linker is invoked. The ld-log wrapper will not capture it, and swapping linkers will not fix it.
  • Incremental relinking — dylib hot reload, -Wl,-r workflows — is where mold feels best interactively, and CI never does it. Developer experience and the CI number point in opposite directions here.

Parallelism contention and filesystem choice

mold spawns threads. If the link runs while compile jobs are still in flight, both fight for the same cores and the measured link time inflates. Isolate the link step before you measure it. Filesystem matters as well: overlayfs and network-backed volumes both hurt mold more than they hurt a mostly single-threaded lld, because parallel I/O exposes latency that a serial linker amortizes away.

Practical Adoption Notes for Mold in CI

Version pinning and keeping the fallback path

Pin the exact linker build in your image and log mold --version in the job. Do not depend on the distro package if you want 3.0 behavior; the distro will lag. Keep an escape hatch:

## fall back without editing the build system
export LINKER_UNDER_TEST=/usr/bin/ld.lld

Because selection lives in an environment variable rather than in CMakeLists.txt, a broken mold release is a one-line CI change, not a commit race.

Failure modes worth testing before you flip the switch

  • Flag gaps. Anything exotic: --gdb-index, sanitizer coverage, --icf variants, custom linker scripts, -Wl,--build-id=sha1.
  • Hardcoded ld lookups. Some toolchains resolve ld by absolute path. The -B prefix trick fails silently there — verify the wrapper log actually has entries.
  • Cross-compilation. Test your target triples; don't assume host support implies cross support.
  • Split debug and patchelf. Post-link tooling that rewrites ELF sections is a common breakage point.
  • Toolchain version mismatches. A mold built against a different libstdc++ than your container is a runtime failure, not a build failure.
💪

Run the wrapper on your existing pipeline for a week before changing anything. The TSV alone tells you whether the linker deserves your attention, and it costs about 2 ms per link.

Where the Mold 3.0 Speedup Actually Lands

Mold 3.0 is a real release following a Rust rewrite, and fast linking is a legitimate goal the project has pursued for years. But the linker is the wrong lever for most CI pipelines, and a release announcement is not evidence about your build.

Here is how I would decide. Run ld-log for a week and sum link time across every invocation in each job. If that sum is under 5% of job wall-clock, do nothing. If it is 5–15% and you are still on GNU ld, switching to mold is a clean, low-risk 5–10% win — and so is switching to lld. If you are already on lld and link is still above 15%, the problem is almost certainly fat LTO, too many test binaries each paying a full link, or debug info you can move out of the default profile. All of those are cheaper fixes than a new linker.

The place mold 3.0 will actually change your day is the developer loop: debug builds, incremental relinks, test binaries. Measure that separately from CI, because the two numbers will not agree, and only one of them is on the critical path for your team's patience.

Further Reading

Share this post

More posts

Comments