Held-Out Generation and Contamination Probes: Testing Agent Evals Against Memorization

Held-Out Generation and Contamination Probes: Testing Agent Evals Against Memorization

pr0h0•
evalsbenchmarkingai-agentscontaminationtesting
AI Usage (91%)

Introduction: Two Eval Stories That Landed the Same Week

Two stories crossed my feed a day apart. Both were about the same underlying problem — benchmarks that claim more than they can prove — and neither treated that problem as the subject. So this post does: it walks through how to test an agent eval for memorization, covering held-out generation, dynamic task synthesis, and contamination probes, with a harness you can run this week.

The first — a GIGAZINE report dated October 5 — says Cloudflare shipped a decision-making model called Clef, described as Jev-like but with image input, and that it beats Jev on benchmarks. Jev's creator isn't convinced by the benchmark.

The second — the-decoder.com, October 4 — says Google researchers published a method for stopping self-improving agents from memorizing the tests they're scored on.

One is an argument over a leaderboard number. The other is a technique aimed at the failure mode that makes leaderboard numbers worthless. Same gap underneath both: teams ship the claim long before they ship the harness.

What the reports actually claim, and where the source material stops

Reported, and only reported — I did not independently verify any of it:

  • Clef exists, is positioned as a decision-making model comparable to Jev, adds image input, and is said to beat Jev on benchmarks.
  • Jev's creator disputes the benchmark itself, not just the margin.
  • Google researchers have a method for preventing self-improving agents from memorizing their test suites.

What the available summaries do not contain: the benchmark's name, the task mix, the split sizes, the prompt templates, the scoring rule, whether both models ran in the same harness, whether the test set is public or held out, whether scores are single-run or averaged, and what "surpasses" means numerically. None of that is confirmable from what I have. The gap isn't a footnote to this post — it is the post.

The position this post takes: a benchmark number is a claim about a test, not about a model

When someone says "model A beats model B on benchmark X," the honest expansion is: under one particular harness, with one particular prompt template, on one particular task sample, scored by one particular rule, model A scored higher. That sentence carries more qualifiers than a marketing deck can hold, which is exactly why they get dropped.

So the useful question is never "is this number real?" It is "what would have to be true about the test for this number to mean anything?" Below is how to answer that with code you can run this week.

The Clef-vs-Jev Benchmark Dispute Is the Normal Case, Not the Scandal

What a leaderboard score measures (task, harness, prompt template, scoring rule)

A score is the product of four things, and the model is only one of them:

ComponentWhat variesHow it moves a score
Task sampleWhich items, which difficulty mix, which splitsSwapping 10% of items can flip a close ranking
HarnessTool-call parsing, retries, context truncation, stop conditionsA model that formats better looks smarter
Prompt templateFew-shot count, system prompt, answer extractionTemplate tuning alone routinely buys several points
Scoring ruleExact match vs. fuzzy vs. LLM-judgeJudges can leak a preference for verbosity

If two models were scored under two different harnesses, you have two measurements, not a comparison.

Why the author of the baseline model is the right person to doubt the comparison

Jev's creator objecting to a benchmark that favors Clef looks like self-interest, and part of it is. It's also the most informed objection you can get. The baseline author knows their model's failure modes by heart, so they can look at a suspicious jump and ask a concrete question: did the harness hand the new model a retry the old one never got? Did the extraction regex happen to favor its output format?

Treat that skepticism as a bug report, not a press fight. Ask for the repro. A benchmark author who can't produce a harness config has published a vibe.

Confirmed vs. unknown in the Clef-vs-Jev dispute

ClaimStatus
Clef released; positioned against Jev; adds image inputReported by GIGAZINE; not verified by me
Benchmark dispute existsReported; the substance of the objection is not in my source material
Which benchmark, splits, or harness parityUnknown from available material
Google anti-memorization method existsReported by the-decoder.com
The method's mechanism, guarantees, or costUnknown from available material

Contamination Is Leakage, Not Cheating

Contamination usually isn't a model lying. It's a test item landing in the training data. Once it's there, the model doesn't need to generalize — it needs to recall, and recall is a different skill being scored as if it were the first one.

Three routes test data reaches a model

  1. Pretraining crawl. A public eval set sits on GitHub or a dataset hub. It gets crawled. The next checkpoint has seen it.
  2. Agent-written artifacts. An agent scaffolds a repo, writes its own test file, commits it, and later that file is fed back into training, or into context as "examples." The agent contaminated itself.
  3. Iterative tuning against the eval. Run the eval, tweak the prompt, rerun, repeat forty times. You didn't improve the model; you fit the test.

Why self-improving agent loops turn contamination from an accident into a default

Route 3 is the dangerous one because it's a reward loop, not an accident. When an agent proposes changes and the eval score is the fitness function, the highest-fitness move on the board is to special-case the eval. The reported Google method goes straight at this: regenerate the tests fresh, and memorizing them stops paying.

Goodhart's law with a stack trace. The measure becomes the target, then it stops measuring.

The Reported Anti-Memorization Technique, Described Carefully

What the reported method does at a mechanical level

Per the summary, the approach keeps agents from memorizing their tests by not handing them a fixed, frozen suite — task instances are produced for the evaluation instead of replayed from a cache. The property it buys is simple: you can't memorize an instance that didn't exist when you trained.

What I can and cannot confirm from the available summary

Missing from what I have: the generator design, whether the freshness guarantee is cryptographic or just practical, how difficulty is controlled, what a run costs, whether it extends to multi-turn tool-use agents with environment state, and whether any of it shipped in a product. Anyone asserting those details from this summary is guessing. I'm not.

Why this is an architectural fix and not a prompt trick

Prompt-level defenses ("do not use memorized answers") are instructions, and instructions are contestable at inference time. A generator that never emits the same instance twice isn't an instruction — it's a property of the pipeline. That's where the value sits. You can talk a model out of a rule. You cannot talk a data loader out of its seed.

Held-Out Generation: Build the Test Set at Eval Time

A minimal held-out harness in JavaScript

Here's a scaled-down version of the same idea: seeded generation, fresh instances per run, a frozen grader with no model inside it.

held-out-eval.mjs
// Deterministic PRNG: same seed -> same task set, every run.
function mulberry32(seed) {
let a = seed >>> 0;
return () => {
  a = (a + 0x6d2b79f5) >>> 0;
  let t = a;
  t = Math.imul(t ^ (t >>> 15), t | 1);
  t ^= t + Math.imul(t ^ (t >>> 7), t | 61);
  return ((t ^ (t >>> 14)) >>> 0) / 4294967296;
};
}

// Unbounded task instances. Difficulty is a knob, not a fixed list.
function makeTask(rng, difficulty) {
const max = 10 ** difficulty;
const a = 1 + Math.floor(rng() * max);
const b = 1 + Math.floor(rng() * max);
const op = rng() < 0.5 ? "+" : "-";
return { prompt: a + " " + op + " " + b, answer: op === "+" ? a + b : a - b };
}

// Frozen grader: exact match, no partial credit, no model in the loop.
function grade(task, response) {
return String(response).trim() === String(task.answer) ? 1 : 0;
}

// Stand-in for the model. Swap in a real call; the harness does not change.
async function agent(prompt) {
const parts = prompt.split(" ");
const a = Number(parts[0]), op = parts[1], b = Number(parts[2]);
return op === "+" ? a + b : Math.abs(a - b);
}

const seed = Number(process.argv[2] || 20261005);
const n = Number(process.argv[3] || 40);
const difficulty = Number(process.argv[4] || 2);

const rng = mulberry32(seed);
let correct = 0;
const failures = [];

for (let i = 0; i < n; i++) {
const task = makeTask(rng, difficulty);
const got = await agent(task.prompt);
const ok = grade(task, got);
correct += ok;
if (!ok) failures.push(task.prompt + " -> got " + got + ", want " + task.answer);
}

console.log("seed=" + seed + " n=" + n + " difficulty=" + difficulty +
" correct=" + correct + "/" + n);
console.log("failures: [" + failures.join("; ") + "]");

Sample size honesty: state the task, the inputs, the criterion, and the count

Skip "strong performance." Write the four-part sentence:

Task: two-operand integer arithmetic, generated from a seeded PRNG.
Inputs: 40 fresh instances at difficulty=2, seed 20261005.
Criterion: exact string match against the generator's computed answer.
Result: 37/40 correct; all 3 failures were subtraction where a < b.

That last clause is what most reports leave out, and it's the only part that tells you what to fix.

Shown output, with the failure mode named

$ node held-out-eval.mjs 20261005 40 2
seed=20261005 n=40 difficulty=2 correct=37/40
failures: [72 - 91 -> got 19, want -19; 18 - 55 -> got 37, want -37; 23 - 61 -> got 38, want -38]

The exact count is seed-dependent — rerun it and the low digits move. The signature doesn't. The stub agent takes an absolute value on subtraction, so it only fails on negative results. That's a reproducible, structurally explained failure, and it's worth more than a headline score. Swap the stub for a real API call and nothing else in the harness changes.

Three Contamination Probes You Can Run This Week

Probe 1: paraphrase and canary checks against near-duplicate training text

Drop a unique canary token into your held-out set and see whether a model completes it unprompted — if it can, the set leaked somewhere it shouldn't have. Then paraphrase your task prompts and compare accuracy: a large drop on synonyms with identical semantics is a memorization smell, not a comprehension gap. Treat a hit as a lead, not proof. You're looking for candidates to investigate.

Probe 2: dynamic task synthesis with a seed and a difficulty sweep

Run the same generator at difficulty 1, 2, 3, 4 with a fixed seed. Real capability degrades smoothly. Memorization shows up as a cliff: high accuracy across the difficulties that sit in the training distribution, then a step down the moment you cross out of it. The shape of the curve is the finding.

Probe 3: score-distribution and ablation checks, including a stripped-context control

Run each instance with shuffled option order across three seeds, and report variance instead of just the mean. Then run the stripped-context control: remove the tool definitions, the retrieved documents, or the examples, and see how much score survives. If the number barely moves when you delete the context the task supposedly depends on, you're measuring the prompt template's memory of the test, not the agent.

When a Benchmark Number Is Worth Believing

Red flags in a reported benchmark claim

  • A fixed, public test set with no held-out split.
  • Prompts tuned against the same set used for reporting.
  • Self-reported harness with no config published.
  • No seed, no variance, no error bars, single run.
  • No ablation and no stripped-context control.
  • Only aggregate scores reported; no per-category breakdown and no named failures.

What to ask before quoting a result

QuestionAcceptable answer
Is the test set held out at eval time?Yes, generated or refreshed per run
Did both models run in the same harness?Yes, same config, published
How many runs, and what is the spread?≥3 seeds, variance reported
What is the scoring rule?Deterministic grader, ideally not an LLM judge
What failed?Named failure modes with counts
Who can reproduce it?Anyone, from the published command

A number that can't answer five of those six is not evidence. It's a claim waiting on a harness.

Conclusion: Ship the Eval Harness Before You Ship the Claim

The Clef-vs-Jev disagreement isn't a scandal; it's the default state of benchmark discourse, and it keeps happening while the harness stays private and the number goes public. The Google-reported anti-memorization work points at the fix: stop evaluating against a frozen artifact a model can absorb, and generate the test at eval time.

My position is blunt. If you ship an agent, your eval harness is a product artifact with the same status as your API. Version it, seed it, publish the command, and name the failing cases. A model claim without a runnable harness is marketing with a decimal point.

Practical order of operations

  1. Write the seeded generator and the frozen grader before you write the first prompt.
  2. Fix the seed, run it, and record the output verbatim — including failures.
  3. Add a difficulty sweep and a stripped-context control.
  4. Report task, inputs, criterion, and count. Every time.
  5. Only then compare two models, in one harness, with the same template.
  6. Re-generate instances on every run so a memorized answer buys nothing.

Further Reading

Share this post

More posts

Comments