
Running Mellum2.1 Locally for CI Repair: Where a 12B MoE Model Holds Up
Introduction: A 12B MoE Coding Model for Local CI Repair
Mellum2.1 is worth running locally for narrow CI repair loops and PR triage. As a general coding agent, it isn't. This post is about where that line sits: what the reports about the 12B MoE open-weight model actually establish, what sparse activation implies for a workload shaped like a failing test and a small diff, and how to find your own answer on your own hardware rather than on someone else's leaderboard.
Mellum2.1 is a ~12B sparse mixture-of-experts open-weight model from JetBrains, reported on 2026-10-08 and 2026-10-09 by three secondary outlets. I have not run it. Not "I haven't used it much" — I have no weights, no model card, and no license terms in front of me, because the supplied material contains none of those. So everything below sits in one of three buckets: what the reports actually say, what the MoE mechanism implies in general, and what you'd have to measure before any of it becomes a fact about your repo.
What the Reports Say About Mellum2.1 — And What They Don't
What Is Confirmed From the Supplied Reports
Three secondary reports — MarkTechPost (2026-10-08), shattered.io (2026-10-09), and tech-insider.org (2026-10-08), all surfaced through Google News RSS — describe a 12B Mixture-of-Experts open coding-agent model from JetBrains with a reported 47% on SWE-Bench. That's the entire confirmed surface: a vendor name, a model name and version, an architecture family, a size figure, a benchmark figure, and a date range.
Note what kind of source these are: aggregator and news write-ups, not a primary vendor bulletin. No JetBrains release page, no model card, no config file, no license. Treat every number below as a report, not a specification.
The 47% SWE-Bench Number Needs a Harness Before It Means Anything
"47% on SWE-Bench" isn't a measurement until you know the harness around it. A bare figure leaves at least these questions open, and this list is inference on my part — the reports don't resolve any of it:
- Which SWE-Bench variant — Lite, Verified, or Full? The subsets aren't comparable.
- pass@1 or pass@k? A pass@10 under best-of sampling is a different claim from a single-shot resolve rate.
- What agent scaffold? Retrieval strategy, file-selection heuristic, and repair-loop policy move the number more than most model changes do.
- What tool budget and retry policy? A harness allowed five self-repair attempts will beat one allowed zero.
- Model alone, or model plus a tuned agent pipeline? The report says "open coding-agent model", which doesn't disambiguate.
I'm not saying the 47% is wrong. I'm saying it's unattributable as written, and an unattributable benchmark number doesn't belong in a decision memo.
Fields the Reports Leave Blank: License, Context Length, Serving Requirements
These are unknown, and I'm deliberately not guessing:
- License terms. Open-weight is not open-source. Commercial use, redistribution, and derivative terms are all unstated here.
- Total vs. active parameters. The reports say 12B. Whether that's total parameters, active parameters per token, or a rounded middle figure isn't stated.
- Context length. No window size given.
- Quantization support. No mention of GGUF, AWQ, or any other format.
- Serving requirements. No VRAM figures, no reference runtime.
- Exact repository or model identifier. Nothing to
pullfrom.
Check all six against the vendor's own release page before you download a single weight file.
Why a 12B MoE Fits CI Repair Better Than a Frontier API
The Shape of the CI Repair Workload
CI repair is short-horizon, high-volume, latency-sensitive, and — critically — it ships with a machine-checkable acceptance test. The job either reruns green or it doesn't. You don't need a model that reasons about system design for twenty minutes; you need one that reads a stack trace, a touched file, and a test file, and emits a small diff. Then execution tells you whether it worked.
Sparse activation is attractive for exactly that shape. In an MoE layer, each token is routed to a subset of experts, so decode cost per token is paid on the active subset rather than the full parameter count. That's the mechanism, and it's standard. Whether it's fast enough on your hardware, at your batch size, with your quantization, is a measurement I haven't taken — mark it untested until you do.
Reproducibility and Data Locality Favor Running Locally
Two unglamorous reasons to prefer local, and both matter more than the benchmark:
- A pinned weight version is reproducible across months. A hosted endpoint behind a moving alias isn't. If you're running an eval suite to decide whether a harness change helped, the model can't silently change underneath you mid-quarter.
- Source logs and private repo slices stay on the runner. CI logs contain test names, environment fragments, and occasionally secrets that leaked into output. Shipping them to a third party is a data-handling decision, not a default.
One caveat that gets lost: a locally-run model is still a third-party artifact. It's untrusted input that produces diffs against your repository. Local doesn't mean trusted; it means you control the blast radius.
Building the CI Repair Loop You Actually Want
The Harness Is the Product, Not the Model
The loop that works is boring on purpose:
failing job → collect the failing test name, the log tail, the touched file, the relevant test file → prompt → patch → apply on a throwaway branch → rerun the specific job → accept only if the target test passes and nothing else regresses.
Two hard caps, both non-negotiable: maximum diff size in lines, and maximum repair iterations. Without them you get a model confidently expanding scope until the budget runs out.
A Minimal Node Runner With Executable Verification
This drives an OpenAI-compatible local endpoint, works in a detached scratch worktree, and never pushes anywhere.
import { execFileSync } from "node:child_process";
const REPO = process.cwd();
const ENDPOINT = "http://127.0.0.1:8000/v1/chat/completions";
const MODEL = process.env.MELLUM_MODEL_ID; // verify the real id from the vendor page
const MAX_ITER = 3;
const MAX_DIFF_LINES = 120;
function run(cmd, args, cwd, allowFail = false) {
try {
return { ok: true, out: execFileSync(cmd, args, { cwd, encoding: "utf8" }) };
} catch (err) {
if (!allowFail) throw err;
return { ok: false, out: (err.stdout || "") + (err.stderr || "") };
}
}
async function propose(context) {
const res = await fetch(ENDPOINT, {
method: "POST",
headers: { "content-type": "application/json" },
body: JSON.stringify({
model: MODEL,
temperature: 0,
messages: [
{ role: "system", content: "Return only a unified diff. No prose." },
{ role: "user", content: context },
],
}),
});
if (!res.ok) throw new Error("endpoint returned " + res.status);
const json = await res.json();
return json.choices[0].message.content;
}
const scratch = mkdtempSync(join(tmpdir(), "ci-repair-"));
run("git", ["worktree", "add", "--detach", scratch, "HEAD"], REPO);
const [cmd, ...args] = process.argv[2].split(" "); // e.g. npx vitest run src/pay.test.ts
const started = Date.now();
for (let i = 1; i <= MAX_ITER; i++) {
const before = run(cmd, args, scratch, true);
if (before.ok) {
console.log(JSON.stringify({ iteration: i, note: "already green, no repair" }));
break; // rerun-before-repair gate
}
const patch = await propose(
"FAILING COMMAND: " + cmd + " " + args.join(" ") + "\n\nLOG TAIL:\n" +
before.out.split("\n").slice(-80).join("\n")
);
const diffLines = patch.split("\n").length;
if (diffLines > MAX_DIFF_LINES) {
console.log(JSON.stringify({ iteration: i, accepted: false, reason: "diff too large", diffLines }));
break;
}
writeFileSync(join(scratch, "candidate.patch"), patch);
const applied = run("git", ["apply", "--whitespace=fix", "candidate.patch"], scratch, true);
if (!applied.ok) {
console.log(JSON.stringify({ iteration: i, accepted: false, reason: "patch did not apply", diffLines }));
continue;
}
const after = run(cmd, args, scratch, true);
console.log(JSON.stringify({
iteration: i,
accepted: after.ok,
targetPassed: after.ok,
diffLines,
seconds: Math.round((Date.now() - started) / 1000),
}));
if (after.ok) break;
}Constraints baked in: --detach worktree so no branch is touched, 127.0.0.1 only, no git push, no git commit, and the loop stops on the first green run or the iteration cap.
Log What You Claim, In a Fixed Record Shape
Record one row per task: task id, iterations, accepted, target test passed, new failures, diff lines, wall-clock seconds. Then report in the style-guide shape — task, inputs, criterion, count.
I didn't run Mellum2.1, because I haven't verified a weight identifier or a license. So the table below is the record format, not a result:
| Task id | Iterations | Accepted | Target test | New failures | Diff lines | Wall clock (s) |
|---|---|---|---|---|---|---|
| rep-001 | not run | not run | not run | not run | not run | not run |
Filling that row with plausible-looking numbers is the single easiest way to make this article worthless.
Where the Active-Parameter Budget Stops Being Enough
Cross-File and Cross-Package Refactors Break the Iteration Loop
Long dependency chains require holding many files in working memory at once. A small per-token active budget is the wrong shape for that, and no iteration budget fixes it. The failure mode to watch for is a patch that's locally correct and globally broken: it updates a call site, misses the interface in another package, and the target test still passes because that test never exercised the other path. Rerunning the specific job won't catch this. Only the full suite does.
Unfamiliar Frameworks and Hallucinated Helper Imports
The model will invent test utilities, config keys, and helper imports that don't exist in your repo. Execution feedback catches the ones that break at import or compile time. It does not catch a helper that exists but behaves differently than assumed. Mitigation: constrain the loop to repos and stacks you can review in under a minute, and never auto-merge.
Flaky Tests and Ambiguous Failures Produce Confident Wrong Patches
If the failure is nondeterministic, or the log tail is incomplete, the loop will confidently patch the wrong thing — and a passing rerun may simply mean the flake didn't fire. Hence the rerun-before-repair gate: if the job goes green on an unmodified tree, exit the loop and patch nothing.
Workflow Files, Secrets, and the CI Trust Boundary
This is a hard limit, not a guideline. The agent must not be permitted to modify .github/workflows/*, release configuration, dependency lockfiles, or anything reachable by a privileged trigger. It must never see repository secrets. Enforce this with a path allowlist in the harness and a CI permission model that gives the repair job read-only tokens — not by asking the model nicely in a system prompt.
Test output, issue text, and commit messages are untrusted content that flows straight into your prompt. A crafted string in a log tail can instruct the model to add a step that exfiltrates environment variables. Keep the repair job's credentials scoped to nothing, and treat any patch that touches workflow or config files as rejected regardless of how green the tests are.
An Adoption Protocol Before You Trust the Benchmark Number
Build the Task Set From Your Own CI History
Take 20–50 real historical CI failures from the repo you care about, and use the original fix commit as ground truth. Four criteria, all required: the target test passes, the full suite has no new failures, the diff is under a fixed line cap, and a human accepts the patch. Anything less and you're measuring the harness's ability to make tests green rather than to fix bugs.
Report Results Honestly: Counts, Not Adjectives
Counts, not adjectives. Ran three prompts? Say three and call it an anecdote — that's a sample size, not a result. Group failures by class so the number points at a fix: wrong file selected, invented API, flaky-test chasing, diff over cap. A 60% acceptance rate with 80% of failures in one class is more actionable than 75% spread evenly.
Serving Checklist: Measure VRAM, Throughput, and Cold Start
Before committing hardware, record: VRAM headroom, tokens/sec at your actual batch size, cold-start time, and quantized-vs-unquantized quality on the same task set. Measure, do not assert:
## terminal 1 — server with verbose logging
./llama-server --model ./weights/model.gguf --n-gpu-layers 99 --port 8000
## terminal 2 — what is actually resident, polled during a run
nvidia-smi --query-gpu=memory.used,memory.total,utilization.gpu --format=csv -l 2
I have no figures to give you here. Any number I invented would be wrong for your card, your context length, and your quantization.
Conclusion — A Narrow Tool, Not a Replacement
The Position: Adopt for CI Repair, Not General Coding
Adopt Mellum2.1 for high-volume, short-horizon CI repair and PR triage where execution feedback exists and a human reviews the diff. Do not adopt it as a general coding agent, and do not quote 47% in a decision memo without the harness details — variant, pass@k, scaffold, and retry policy. The MoE sparsity trade-off is what makes it survive the first workload and fail the second: paying decode cost on the active subset is cheap when the task is one file and one test, and irrelevant when the task is a cross-package refactor where the bottleneck is context, not throughput.
Further Reading
- JetBrains Releases Mellum2.1: A 12B MoE Open Model for Coding Agents — secondary report, MarkTechPost, 2026-10-08.
- JetBrains Mellum2.1: 12B Model Hits 47% SWE-Bench — secondary report, shattered.io, 2026-10-09.
- JetBrains Mellum 2.1: 12B MoE Model Targets Coding Agents — secondary report, tech-insider.org, 2026-10-08.
- SWE-bench: Can Language Models Resolve Real-World GitHub Issues? — the benchmark paper, useful for understanding why the Lite, Verified, and Full subsets are not interchangeable.
No primary JetBrains release page or model card was available in the supplied material. Verify license terms, the exact model identifier, and the benchmark harness against the vendor's own page before you download or deploy anything.


