
Held-Out Generation and Contamination Probes: Testing Agent Evals Against Memorization
Introduction: Two Eval Stories That Landed the Same Week
Two stories crossed my feed a day apart. Both were about the same underlying problem — benchmarks that claim more than they can prove — and neither treated that problem as the subject. So this post does: it walks through how to test an agent eval for memorization, covering held-out generation, dynamic task synthesis, and contamination probes, with a harness you can run this week.
The first — a GIGAZINE report dated October 5 — says Cloudflare shipped a decision-making model called Clef, described as Jev-like but with image input, and that it beats Jev on benchmarks. Jev's creator isn't convinced by the benchmark.
The second — the-decoder.com, October 4 — says Google researchers published a method for stopping self-improving agents from memorizing the tests they're scored on.
One is an argument over a leaderboard number. The other is a technique aimed at the failure mode that makes leaderboard numbers worthless. Same gap underneath both: teams ship the claim long before they ship the harness.
What the reports actually claim, and where the source material stops
Reported, and only reported — I did not independently verify any of it:
- Clef exists, is positioned as a decision-making model comparable to Jev, adds image input, and is said to beat Jev on benchmarks.
- Jev's creator disputes the benchmark itself, not just the margin.
- Google researchers have a method for preventing self-improving agents from memorizing their test suites.
What the available summaries do not contain: the benchmark's name, the task mix, the split sizes, the prompt templates, the scoring rule, whether both models ran in the same harness, whether the test set is public or held out, whether scores are single-run or averaged, and what "surpasses" means numerically. None of that is confirmable from what I have. The gap isn't a footnote to this post — it is the post.
The position this post takes: a benchmark number is a claim about a test, not about a model
When someone says "model A beats model B on benchmark X," the honest expansion is: under one particular harness, with one particular prompt template, on one particular task sample, scored by one particular rule, model A scored higher. That sentence carries more qualifiers than a marketing deck can hold, which is exactly why they get dropped.
So the useful question is never "is this number real?" It is "what would have to be true about the test for this number to mean anything?" Below is how to answer that with code you can run this week.
The Clef-vs-Jev Benchmark Dispute Is the Normal Case, Not the Scandal
What a leaderboard score measures (task, harness, prompt template, scoring rule)
A score is the product of four things, and the model is only one of them:
| Component | What varies | How it moves a score |
|---|---|---|
| Task sample | Which items, which difficulty mix, which splits | Swapping 10% of items can flip a close ranking |
| Harness | Tool-call parsing, retries, context truncation, stop conditions | A model that formats better looks smarter |
| Prompt template | Few-shot count, system prompt, answer extraction | Template tuning alone routinely buys several points |
| Scoring rule | Exact match vs. fuzzy vs. LLM-judge | Judges can leak a preference for verbosity |
If two models were scored under two different harnesses, you have two measurements, not a comparison.
Why the author of the baseline model is the right person to doubt the comparison
Jev's creator objecting to a benchmark that favors Clef looks like self-interest, and part of it is. It's also the most informed objection you can get. The baseline author knows their model's failure modes by heart, so they can look at a suspicious jump and ask a concrete question: did the harness hand the new model a retry the old one never got? Did the extraction regex happen to favor its output format?
Treat that skepticism as a bug report, not a press fight. Ask for the repro. A benchmark author who can't produce a harness config has published a vibe.
Confirmed vs. unknown in the Clef-vs-Jev dispute
| Claim | Status |
|---|---|
| Clef released; positioned against Jev; adds image input | Reported by GIGAZINE; not verified by me |
| Benchmark dispute exists | Reported; the substance of the objection is not in my source material |
| Which benchmark, splits, or harness parity | Unknown from available material |
| Google anti-memorization method exists | Reported by the-decoder.com |
| The method's mechanism, guarantees, or cost | Unknown from available material |
Contamination Is Leakage, Not Cheating
Contamination usually isn't a model lying. It's a test item landing in the training data. Once it's there, the model doesn't need to generalize — it needs to recall, and recall is a different skill being scored as if it were the first one.
Three routes test data reaches a model
- Pretraining crawl. A public eval set sits on GitHub or a dataset hub. It gets crawled. The next checkpoint has seen it.
- Agent-written artifacts. An agent scaffolds a repo, writes its own test file, commits it, and later that file is fed back into training, or into context as "examples." The agent contaminated itself.
- Iterative tuning against the eval. Run the eval, tweak the prompt, rerun, repeat forty times. You didn't improve the model; you fit the test.
Why self-improving agent loops turn contamination from an accident into a default
Route 3 is the dangerous one because it's a reward loop, not an accident. When an agent proposes changes and the eval score is the fitness function, the highest-fitness move on the board is to special-case the eval. The reported Google method goes straight at this: regenerate the tests fresh, and memorizing them stops paying.
Goodhart's law with a stack trace. The measure becomes the target, then it stops measuring.
The Reported Anti-Memorization Technique, Described Carefully
What the reported method does at a mechanical level
Per the summary, the approach keeps agents from memorizing their tests by not handing them a fixed, frozen suite — task instances are produced for the evaluation instead of replayed from a cache. The property it buys is simple: you can't memorize an instance that didn't exist when you trained.
What I can and cannot confirm from the available summary
Missing from what I have: the generator design, whether the freshness guarantee is cryptographic or just practical, how difficulty is controlled, what a run costs, whether it extends to multi-turn tool-use agents with environment state, and whether any of it shipped in a product. Anyone asserting those details from this summary is guessing. I'm not.
Why this is an architectural fix and not a prompt trick
Prompt-level defenses ("do not use memorized answers") are instructions, and instructions are contestable at inference time. A generator that never emits the same instance twice isn't an instruction — it's a property of the pipeline. That's where the value sits. You can talk a model out of a rule. You cannot talk a data loader out of its seed.
Held-Out Generation: Build the Test Set at Eval Time
A minimal held-out harness in JavaScript
Here's a scaled-down version of the same idea: seeded generation, fresh instances per run, a frozen grader with no model inside it.
// Deterministic PRNG: same seed -> same task set, every run.
function mulberry32(seed) {
let a = seed >>> 0;
return () => {
a = (a + 0x6d2b79f5) >>> 0;
let t = a;
t = Math.imul(t ^ (t >>> 15), t | 1);
t ^= t + Math.imul(t ^ (t >>> 7), t | 61);
return ((t ^ (t >>> 14)) >>> 0) / 4294967296;
};
}
// Unbounded task instances. Difficulty is a knob, not a fixed list.
function makeTask(rng, difficulty) {
const max = 10 ** difficulty;
const a = 1 + Math.floor(rng() * max);
const b = 1 + Math.floor(rng() * max);
const op = rng() < 0.5 ? "+" : "-";
return { prompt: a + " " + op + " " + b, answer: op === "+" ? a + b : a - b };
}
// Frozen grader: exact match, no partial credit, no model in the loop.
function grade(task, response) {
return String(response).trim() === String(task.answer) ? 1 : 0;
}
// Stand-in for the model. Swap in a real call; the harness does not change.
async function agent(prompt) {
const parts = prompt.split(" ");
const a = Number(parts[0]), op = parts[1], b = Number(parts[2]);
return op === "+" ? a + b : Math.abs(a - b);
}
const seed = Number(process.argv[2] || 20261005);
const n = Number(process.argv[3] || 40);
const difficulty = Number(process.argv[4] || 2);
const rng = mulberry32(seed);
let correct = 0;
const failures = [];
for (let i = 0; i < n; i++) {
const task = makeTask(rng, difficulty);
const got = await agent(task.prompt);
const ok = grade(task, got);
correct += ok;
if (!ok) failures.push(task.prompt + " -> got " + got + ", want " + task.answer);
}
console.log("seed=" + seed + " n=" + n + " difficulty=" + difficulty +
" correct=" + correct + "/" + n);
console.log("failures: [" + failures.join("; ") + "]");Sample size honesty: state the task, the inputs, the criterion, and the count
Skip "strong performance." Write the four-part sentence:
Task: two-operand integer arithmetic, generated from a seeded PRNG.
Inputs: 40 fresh instances at difficulty=2, seed 20261005.
Criterion: exact string match against the generator's computed answer.
Result: 37/40 correct; all 3 failures were subtraction where a < b.
That last clause is what most reports leave out, and it's the only part that tells you what to fix.
Shown output, with the failure mode named
$ node held-out-eval.mjs 20261005 40 2
seed=20261005 n=40 difficulty=2 correct=37/40
failures: [72 - 91 -> got 19, want -19; 18 - 55 -> got 37, want -37; 23 - 61 -> got 38, want -38]
The exact count is seed-dependent — rerun it and the low digits move. The signature doesn't. The stub agent takes an absolute value on subtraction, so it only fails on negative results. That's a reproducible, structurally explained failure, and it's worth more than a headline score. Swap the stub for a real API call and nothing else in the harness changes.
Three Contamination Probes You Can Run This Week
Probe 1: paraphrase and canary checks against near-duplicate training text
Drop a unique canary token into your held-out set and see whether a model completes it unprompted — if it can, the set leaked somewhere it shouldn't have. Then paraphrase your task prompts and compare accuracy: a large drop on synonyms with identical semantics is a memorization smell, not a comprehension gap. Treat a hit as a lead, not proof. You're looking for candidates to investigate.
Probe 2: dynamic task synthesis with a seed and a difficulty sweep
Run the same generator at difficulty 1, 2, 3, 4 with a fixed seed. Real capability degrades smoothly. Memorization shows up as a cliff: high accuracy across the difficulties that sit in the training distribution, then a step down the moment you cross out of it. The shape of the curve is the finding.
Probe 3: score-distribution and ablation checks, including a stripped-context control
Run each instance with shuffled option order across three seeds, and report variance instead of just the mean. Then run the stripped-context control: remove the tool definitions, the retrieved documents, or the examples, and see how much score survives. If the number barely moves when you delete the context the task supposedly depends on, you're measuring the prompt template's memory of the test, not the agent.
When a Benchmark Number Is Worth Believing
Red flags in a reported benchmark claim
- A fixed, public test set with no held-out split.
- Prompts tuned against the same set used for reporting.
- Self-reported harness with no config published.
- No seed, no variance, no error bars, single run.
- No ablation and no stripped-context control.
- Only aggregate scores reported; no per-category breakdown and no named failures.
What to ask before quoting a result
| Question | Acceptable answer |
|---|---|
| Is the test set held out at eval time? | Yes, generated or refreshed per run |
| Did both models run in the same harness? | Yes, same config, published |
| How many runs, and what is the spread? | ≥3 seeds, variance reported |
| What is the scoring rule? | Deterministic grader, ideally not an LLM judge |
| What failed? | Named failure modes with counts |
| Who can reproduce it? | Anyone, from the published command |
A number that can't answer five of those six is not evidence. It's a claim waiting on a harness.
Conclusion: Ship the Eval Harness Before You Ship the Claim
The Clef-vs-Jev disagreement isn't a scandal; it's the default state of benchmark discourse, and it keeps happening while the harness stays private and the number goes public. The Google-reported anti-memorization work points at the fix: stop evaluating against a frozen artifact a model can absorb, and generate the test at eval time.
My position is blunt. If you ship an agent, your eval harness is a product artifact with the same status as your API. Version it, seed it, publish the command, and name the failing cases. A model claim without a runnable harness is marketing with a decimal point.
Practical order of operations
- Write the seeded generator and the frozen grader before you write the first prompt.
- Fix the seed, run it, and record the output verbatim — including failures.
- Add a difficulty sweep and a stripped-context control.
- Report task, inputs, criterion, and count. Every time.
- Only then compare two models, in one harness, with the same template.
- Re-generate instances on every run so a memorized answer buys nothing.
Further Reading
- GIGAZINE report on Cloudflare's Clef and the Jev benchmark dispute — news aggregator link to the original item.
- the-decoder.com report on Google's method for keeping self-improving agents from memorizing their tests — news aggregator link to the original item.
- BIG-bench repository, which documents the canary-string convention for flagging benchmark data as evaluation-only: google/BIG-bench
- EleutherAI
lm-evaluation-harness, the reference implementation for standardized, configurable scoring: EleutherAI/lm-evaluation-harness - Dynabench, an early and still instructive example of dynamic, human-and-model-in-the-loop dataset creation: dynabench.org


