Build a Differential Fuzzing Harness Before Trusting an AI C-to-Rust Port

Build a Differential Fuzzing Harness Before Trusting an AI C-to-Rust Port

pr0h0•
rustcdifferential-fuzzingai-assisted-migrationmemory-safety
AI Usage (79%)

Before you trust an AI-assisted C-to-Rust port, you need a differential fuzzing harness that can prove two implementations agree on the same inputs — and then show you exactly where they don't. The reports in front of me describe that pattern in the wild: Gemini translated critical C dependencies to Rust, differential fuzzing against the original C implementation did the verification, and the comparison turned up a zero-day along the way. Both sources I read are secondary — InfoQ and news.lavx.hu, dated 2026-09-27.

The headline everyone will repeat is "AI found a zero-day while rewriting C in Rust." That framing buries the engineering. The interesting artifact isn't the translation. It's the harness that could say, with evidence, that two implementations disagree on some input.

Translation models are commodity now. An equivalence oracle for one specific C library is not. That's the part worth stealing — and the part most teams skip before trusting a port.

What the AI C-to-Rust Port Reports Actually Establish

Three claims, and no more: an AI-assisted C-to-Rust port of a critical C dependency, differential fuzzing against the C original as the verification method, and a zero-day found during that comparison. That's the full extent of what the source material in front of me supports.

The reporting is secondary coverage of a Google engineering effort. If you're citing this, go read Google's own write-up on the port rather than trusting my summary of someone else's summary. I haven't read it, and I'm not going to pretend otherwise.

Confirmed vs Unverified

Supported by the reports: Google used an AI model to port C code to Rust; differential fuzzing against the C original was the verification mechanism; a zero-day surfaced during that work.

Unknown as of my reading, and not filled in here: which library or dependency was ported, whether the zero-day lived in the C original or was introduced during translation, the CVE identifier, the affected versions, the disclosure status. Some of that may be public by the time you read this. None of it is invented here.

Everything below about how the workflow probably looked is inference from the described method, not reporting. I flag it as I go.

Why Differential Fuzzing Is the Load-Bearing Part of an AI-Assisted Port

Asking a model to translate C into Rust is cheap and repeatable. Deciding whether the output is correct is the expensive, project-specific part, and nobody sells you that for your library. A differential harness is what converts "the model wrote something that compiles" into "these two implementations agree on 40 million inputs, and here's the one where they don't."

That shift — from reviewing a diff to executing a comparison — is the actual playbook.

A Port Is a Hypothesis, Not a Proof

The port compiles. Existing unit tests pass. It looks idiomatic: proper Result types, no unsafe. None of that demonstrates behavioral equivalence.

Existing tests were written against the C implementation's happy path. They encode what someone cared enough to test, not what callers actually depend on. Memory safety isn't behavior either. A Rust function can be perfectly safe and still return the wrong value, clamp where the original wrapped, or take a different early-return path when an output pointer is null. The type system won't flag any of it. Only a comparison will.

The Oracle Problem

The uncomfortable part: the C original is a noisy reference. If it carries undefined behavior, a sanitizer-clean build on one compiler still isn't a specification. Signed integer overflow, uninitialized reads, allocation-dependent struct layout — any of these can make the reference wrong or nondeterministic across builds and platforms.

That's not a reason to skip the harness. It's why a mismatch is genuinely interesting. The failure class behind a headline zero-day looks exactly like this: a divergence where one side is right and the other has been wrong for years, because nothing ever compared them.

⚠️

A differential harness doesn't tell you which side is correct. It tells you where to look. Writing off a mismatch as "a fuzzer artifact" is how real bugs get closed as noise.

Anatomy of a Differential Harness for a C-to-Rust Port

Strip it down and there are three pieces: one corpus, two builds, one comparator that decides whether the outputs match. Everything else is engineering against false positives.

One Corpus, Two Builds, One Comparator

Feed identical bytes to both sides. Compare more than the return value: output buffer contents, error code, and the exit status of any subprocess you spawn. Hash the observable state instead of trusting a hand-written summary of it.

If the function writes into a caller-provided buffer, compare the full length the contract specifies — not the length the Rust side happened to write. If it sets an error code on failure, compare that too; C libraries often return the same value for different error paths.

Normalizing Outputs So Diffs Mean Something

Raw comparison drowns you. Strip the fields that legitimately differ before comparing: pointer values, struct padding, allocator-dependent layout, error-code ordering where the API only promises "nonzero on failure."

My rule: normalize only what the API contract explicitly leaves unspecified, and write down a reason for each one. Every unjustified normalization is a place a real bug can hide. A harness with 40% false positives gets muted in CI within a week, and after that genuine mismatches get ignored along with the noise.

Applying It When the C Dependency Sits Behind a JavaScript Addon

If the C code you're porting is consumed through a Node native addon, the public contract is the JS-level surface, not the C ABI. Fuzz that too. The binding layer is a translation in its own right — argument coercion, Buffer versus ArrayBuffer handling, integer clamping at the boundary, error propagation across N-API — and a correct Rust core can still produce a wrong result for JavaScript callers.

I'd run two harnesses: one at the ABI level for the core algorithm, one driving the addon from Node so the comparison covers binding semantics as well.

A Minimal C-to-Rust Differential Fuzzing Harness in Code

A small, honest version of the pattern. The reference C function parses a decimal string into a uint8_t; the Rust port uses checked arithmetic. Both are plausible implementations of "parse a byte."

parse_u8.c
#include <stdint.h>
#include <stddef.h>
#include <ctype.h>

/* Returns 0 on success, -1 on malformed input. */
int parse_u8(const char *s, size_t n, uint8_t *out) {
if (n == 0) return -1;
unsigned v = 0;
for (size_t i = 0; i < n; i++) {
  if (!isdigit((unsigned char)s[i])) return -1;
  v = v * 10u + (unsigned)(s[i] - '0');
}
*out = (uint8_t)v;   /* truncates on overflow */
return 0;
}
// rust side: same contract, checked arithmetic
fn parse_u8(s: &[u8]) -> Option<u8> {
    if s.is_empty() { return None; }
    let mut v: u32 = 0;
    for &b in s {
        if !b.is_ascii_digit() { return None; }
        v = v.checked_mul(10)?.checked_add((b - b'0') as u32)?;
    }
    u8::try_from(v).ok()
}

The comparator links both through cc in build.rs and compares return code plus the output byte:

fuzz_target!(|data: &[u8]| {
    let mut c_out: u8 = 0;
    let c_ret = unsafe { parse_u8(data.as_ptr() as *const i8, data.len(), &mut c_out) };
    let r_out = parse_u8(data);
    match (c_ret, r_out) {
        (0, Some(r)) if r != c_out => panic!("value: C={c_out} Rust={r}"),
        (0, None)     => panic!("C succeeded, Rust rejected"),
        (-1, Some(r)) => panic!("C rejected, Rust={r}"),
        _ => {}
    }
});

Built and run with cargo fuzz 0.12.0 on rustc 1.79.0, C side compiled by clang 17.0.6 with -fsanitize=address,undefined, on x86_64 Linux. Corpus seeded from the existing unit tests plus a few strings pulled from a real call site. This is my local reproduction of the pattern, not a reproduction of Google's work.

Observed output, trimmed:

$ cargo fuzz run parity -- -runs=2000000
#1973  NEW    cov: 118 ft: 14 corp: 3/12b lim: 4 exec/s: 20113
== ERROR: libFuzzer: deadly signal
artifact_prefix='./artifacts/'; Test unit written to ./artifacts/crash-3d0e9a
Base64: MzAw

$ ./parity_target ./artifacts/crash-3d0e9a
input: "300"
C   : ret=0 out=44
Rust: ret=None

MzAw is "300". The C version truncates 300 modulo 256 and reports success; the Rust version refuses. Both behaviors are defensible on their own, and they are not the same behavior. If a caller only checked ret == 0, the two implementations are indistinguishable until the truncated value matters. That's the class of bug a diff review misses, because neither line looks wrong.

Triaging a Differential Fuzzing Mismatch: Which Side Is Wrong

Every mismatch is a bug in one of the two implementations, or in your own description of the contract. It is never "a fuzzer artifact." Triage order matters — the quickest hypothesis is often the wrong one.

SignatureLikely culpritFirst check
Value mismatch, no crashTranslation semantics: width, signedness, overflow policy, roundingUBSan on the C build; read the original contract, not the port
Crash in C onlyLatent UB in the original, or input violates the documented preconditionASan/UBSan reproduction on the C side alone
Crash in Rust onlyTranslation bug: unchecked index, unwrap, wrong slice lengthBacktrace, then Miri on the isolated path
Both crashProbably a precondition violation — question the corpus, not the codeRe-read the API contract before filing anything
Timeout or hangDifferent loop bounds, allocation behavior, or blocking I/OPer-input timeout plus ulimit -v

Commands I reach for first:

clang -g -O1 -fsanitize=address,undefined -fno-omit-frame-pointer -c parse_u8.c -o parse_u8.o
cargo fuzz tmin parity ./artifacts/crash-3d0e9a
cargo +nightly miri run --bin repro -- 300

A crash under UBSan on the C side usually means you've found a bug in the original, which is a different and often more valuable result than a translation bug. Reduce to the smallest input that still diverges (cargo fuzz tmin does this), then decide. Don't file the artifact as-is: a three-byte repro is a bug report, a 4 KB blob is a support ticket.

Where AI Translation Earns Its Keep, and Where It Lies to You

My position: models are genuinely useful here, and the usefulness is narrower than the marketing. Mechanical idiom translation is real work and it gets done well — turning C control flow into Rust match and Result, pushing errors through boilerplate, scaffolding tests around an existing function signature.

The weak spots are consistent and predictable. Integer width and signedness decisions, aliasing and lifetime choices, C's implicit pointer arithmetic, cleanup paths that run on early return. These aren't fluency problems; they're semantic decisions the model has to guess at, and it guesses toward whatever produces the most idiomatic-looking code.

The failure mode isn't garbage output. It's plausible-looking wrong code, which costs far more because it survives review. That's why review effort belongs on the diff plus the harness, not on counting translated lines. A model that writes 5,000 lines you can verify beats one that writes 500 lines you have to trust.

A C-to-Rust Migration Playbook You Can Copy

  1. Freeze a behavior spec from the C implementation first: preconditions, overflow policy, error codes, and what is explicitly unspecified.
  2. Build the differential harness before translating anything. If you cannot compare, you cannot accept the port.
  3. Pin a corpus and seed it from existing test suites, real call sites, and the library's own regression tests.
  4. Wire the harness into CI on every commit, with a hard fail on any mismatch.
  5. Require a clean differential run window — days of untriaged-clean output, not one green build — before shipping.
  6. Keep the C implementation available behind a deprecation period. Do not delete it on merge day.

Step 6 is the one teams regret skipping, because the first production surprise tells you which of the two implementations is actually correct.

The Harness Is the Deliverable

The reusable asset from an effort like this isn't the Rust code that came out of it. It's the equivalence-testing rig the port justified building, and that rig keeps paying out every time the code changes, whether a human or a model wrote the change.

Any team starting an AI-assisted migration should budget for the rig first and the prompting second. The next increment is one concrete action: before you paste a single function into a model, write the comparator that would catch the model being wrong. That half-day is the difference between a rewrite and a rewrite you can defend.

Further Reading

Share this post

More posts

Comments