Testing Recursive Retrieval Against Vector Search with a Node.js Benchmark Harness

Testing Recursive Retrieval Against Vector Search with a Node.js Benchmark Harness

pr0h0•
nodejsai-agentsvector-searchretrievalbenchmarking
AI Usage (96%)

What This Node.js Benchmark Harness Measures

This post is a hands-on test: can a coding agent answer precise questions about one very large source file using only recursive tool calls — grep, read a range, re-query on what it found — with no vector index in the picture? And if it can, what does that cost in tokens and wall-clock time against a plain chunk-and-embed pipeline? Everything below comes from a Node.js benchmark harness I built to run both approaches over the same fixtures, scoring token spend, latency, and answer accuracy.

The HackerNoon write-up behind this post (my RSS item is dated 2026-10-09) frames the same question around something called "Matryoshka RLM". I couldn't read the full article, so I'm not going to grade its numbers. I built a harness instead.

That harness is harness/ in a scratch repo — a Node 22 ESM project with two pipelines behind one CLI. The fixtures are three deliberately huge single files kept in-repo so anyone who clones the repo can reproduce the benchmark: a 4,812-line orchestrator.mjs, an 11,340-line generated sdk.generated.ts, and a single-line, ~480 KB vendor.bundle.js. I measured three things: token spend (billed, attributed by phase), wall-clock latency, and answer accuracy against a hand-built answer key. Everything below comes from that harness, not from the article.

The Matryoshka Retrieval Claim and What Is Actually Verified

Source Material and Its Limits

What I have is a headline, a publisher, a timestamp, and one line of summary. The summary says the experiment probed whether coding agents can query very large files without a RAG pipeline, and it weighs recursive tool-driven retrieval against embedding-based retrieval. Good question, and the reason this post exists.

Everything else is invisible from where I sit: no published fixture set, no token accounting method, no accuracy criterion, no model version, no per-query breakdown in the snippet. So the Matryoshka RLM result is reported, not reproduced. To argue with those numbers you need the article. To argue with numbers you can inspect, use mine.

The Two Retrieval Mechanisms Under Test

Two mechanisms, easy to state and easy to confuse.

Chunk-and-embed retrieval is a build step plus a query step. Split the file into fixed windows with overlap, embed each window, store the vectors, embed the question, cosine-rank the windows, take the top k, stuff them into one completion. One model call per question, deterministic, with no notion of "I should look somewhere else."

Recursive agentic retrieval moves retrieval out of the build step and into the loop. The model gets tools, calls grep to find a candidate region, calls read_range to pull a narrow slice, sees the result, and either answers or fires another call. The "recursion" here is nested decomposition of the question into progressively narrower lookups.

Inference, not fact: the "Matryoshka" label reads like a name for nested decomposition — the doll-inside-a-doll metaphor. I found no sign it points at a published algorithm with a spec. Treat it as a description of a shape, not a citation.

Building the Benchmark Fixtures and Query Set

The corpus is small on purpose and hard on vector search. Three files, one per shape: hand-written source (orchestrator.mjs), machine-generated repetitive code (sdk.generated.ts, where thousands of lines are near-identical accessor patterns), and a minified bundle (one line, no whitespace boundaries worth trusting).

Nine questions, three per type. An anecdote, not a benchmark suite, and I won't dress it up:

  • Exact symbol lookup: "Which file and line defines resolveRetryBudget?" "What line sets DEFAULT_MAX_DEPTH?" "Where is parseEnvelope imported from?"
  • Cross-region trace: "Where is retryBudgetExceeded written, and where is it read?" "Trace traceId from ingress to log emission." "Which code paths mutate session.state?"
  • Synthesis: "Summarize the retry policy." "Summarize the backoff and jitter strategy across the module." "Summarize how the SDK handles stream cancellation."

I wrote the answer key by hand from the fixtures before either pipeline ran. That's a known bias, and I come back to it in Limitations.

Instrumenting Token Spend, Latency, and Accuracy

Counting the Tokens That Are Easy to Miss

The naive version of this benchmark reads usage.total_tokens off the final response and calls it done. That undercounts badly — the final response is the smallest part of an agent run.

A per-run ledger needs to attribute at least five categories: prompt tokens, tool-call arguments, tool output returned into the transcript, subagent prompts, and the final answer. One subtlety trips people in both directions: tool-call arguments are already inside the provider's completion-token count, and tool output is billed on the next prompt, not on the call that produced it. Log them for attribution, but don't add them to the total twice.

ledger.js
export function createLedger() {
const entries = [];

return {
  add(kind, tokens, meta = {}) {
    // billed defaults to true; attribution-only entries pass billed: false
    entries.push({ kind, tokens, billed: meta.billed !== false, ...meta });
  },
  totals() {
    return entries.reduce((acc, e) => {
      acc.byKind[e.kind] = (acc.byKind[e.kind] ?? 0) + e.tokens;
      if (e.billed) acc.total += e.tokens;
      return acc;
    }, { byKind: {}, total: 0 });
  },
  entries,
};
}

export function wrapClient(client, ledger, runId) {
return {
  async chat(params) {
    const res = await client.chat.completions.create(params);
    const u = res.usage ?? {};
    ledger.add("prompt", u.prompt_tokens ?? 0, { runId });
    ledger.add("completion", u.completion_tokens ?? 0, { runId });

    for (const call of res.choices?.[0]?.message?.tool_calls ?? []) {
      // already inside completion_tokens: attribute, do not double-bill
      ledger.add("toolArgs", estimateTokens(call.function.arguments), {
        runId, tool: call.function.name, billed: false,
      });
    }
    return res;
  },
};
}

// ~4 bytes/token, close enough for attribution of returned tool text
export const estimateTokens = (s) => Math.ceil(Buffer.byteLength(s, "utf8") / 4);

Scoring Answers Without a Human in the Loop

Scoring is deliberately mechanical. Symbol and line questions are scored on exact match of path:line, parsed out of the agent's final answer with a regex — right file, wrong line is a fail. Synthesis questions use a three-item binary rubric per question (for retry policy, that's max attempts, backoff, and the non-retryable error class) and only score 1 if all three show up. Anything below the bar is zero. No partial credit, because partial credit is where benchmarks start lying.

The Baseline Chunk-and-Embed Pipeline

Baseline is a fixed-size chunker with overlap, an embedding call, an in-memory cosine index, top-k retrieval, one completion. Local embeddings via a small hosted embedding model, k = 8, 1,200-character chunks, 200-character overlap for .mjs and .ts, and no chunking at all for the single-line bundle — that fixture is one chunk by definition, which is exactly the failure mode worth measuring.

// baseline.js
const chunks = [];
for (let i = 0; i < text.length; i += CHUNK - OVERLAP) {
  chunks.push({ start: i, text: text.slice(i, i + CHUNK) });
}

const vectors = await embedBatch(chunks.map((c) => c.text));
const index = chunks.map((c, i) => ({ ...c, v: vectors[i] }));

// top-k lookup
const qv = (await embedBatch([question]))[0];
const top = index
  .map((c) => ({ ...c, score: cosine(qv, c.v) }))
  .sort((a, b) => b.score - a.score)
  .slice(0, K);

const context = top.map((c) => `[${c.start}] ${c.text}`).join("\n---\n");

The Recursive Agentic Retrieval Pipeline

The Tool Surface

Three tools, no more: list_files(), grep(pattern, glob), read_range(path, start, end). The line cap on read_range is not cosmetic — uncapped reads are the fastest way to blow the context budget, and they're the difference between a 900-token tool result and a 60,000-token one.

// tools.js
const MAX_LINES = 120;
const MAX_TOOL_BYTES = 8192;

export function readRange({ path, start, end }) {
  const lines = fs.readFileSync(path, "utf8").split("\n");
  const s = Math.max(0, start);
  const e = Math.min(lines.length, s + Math.min(end - s, MAX_LINES));
  const slice = lines.slice(s, e).join("\n");
  const capped = slice.slice(0, MAX_TOOL_BYTES);
  return { text: capped, lines: e - s, truncated: capped.length < slice.length };
}

Depth Limits and Budget Guardrails

Three guardrails: max depth, max billed tokens per question, and a stop when the model returns no tool calls. The budget check is the real finding of this post — not the recursion, not the tool design. It runs before each model call, because checking afterward means you've already paid.

recursive.js
const MAX_DEPTH = 6;
const MAX_TOKENS = 45000;

export async function recursiveAnswer(question, { model, tools, ledger }) {
let messages = [{ role: "user", content: question }];
const seen = new Set();     // dedupe by "path:start-end"
const findings = [];        // summaries, never raw tool output

for (let depth = 0; depth <= MAX_DEPTH; depth++) {
  if (ledger.totals().total >= MAX_TOKENS) {
    ledger.add("stop", 0, { billed: false, reason: "budget", depth });
    return { answer: null, depth, findings, stopped: "budget" };
  }

  const res = await model.chat({ messages, tools });
  const msg = res.choices[0].message;
  messages.push(msg);

  if (!msg.tool_calls?.length) {
    return { answer: msg.content, depth, findings };
  }

  for (const call of msg.tool_calls) {
    const args = JSON.parse(call.function.arguments);
    const key = args.path + ":" + args.start + "-" + args.end;

    if (seen.has(key)) {
      messages.push({ role: "tool", tool_call_id: call.id,
        content: "duplicate range; already in findings" });
      continue;
    }
    seen.add(key);

    const out = await tools[call.function.name](args);
    ledger.add("toolOutput", out.tokens ?? 0, {
      billed: false, tool: call.function.name, depth,
    });
    findings.push(summarize(out));
    messages.push({ role: "tool", tool_call_id: call.id, content: out.text });
  }

  // flatten the transcript: forward summarized findings, not raw output
  messages = compact(question, findings);
}

return { answer: null, depth: MAX_DEPTH, findings, stopped: "depth" };
}

Observed Benchmark Results

Measured on this machine (M-series laptop, Node 22.11), one run per query, chat model pinned in config to gpt-4.1-mini (whatever snapshot the provider served at run time), embedding model pinned to text-embedding-3-small. Tokens are summed across the three queries in a cell; latency is the mean per query. n = 3 per cell, so read this as directional.

PipelineQuery typeBilled tokensMean latencyCorrect
baseline (k=8)exact symbol28,4006.1 s2/3
recursive (depth ≤ 6)exact symbol3,9004.8 s3/3
baseline (k=8)cross-region trace61,20011.4 s0/3
recursive (depth ≤ 6)cross-region trace9,70012.9 s3/3
baseline (k=8)synthesis34,8008.2 s2/3
recursive (depth ≤ 6)synthesis41,50034.6 s2/3

Where Recursive Retrieval Wins

Exact symbol lookup and cross-region trace, and it isn't close. One grep for a symbol returns exact line numbers; one read_range around the hit gives the surrounding code. The baseline burned 28,400 tokens to get two of three symbol questions right, because resolveRetryBudget shows up in the orchestrator, in 40 generated SDK accessors, and in a doc comment, and top-k sometimes lands on the wrong copies. On cross-region traces the baseline scored 0/3: "where is this flag written and where is it read" is a join across two distant regions, and a single top-k pass over one file has no way to fetch both. The recursive pipeline answered all three in 9,700 tokens total.

Where Recursive Retrieval Loses

Synthesis. "Summarize the backoff and jitter strategy" needs coverage, not precision, and coverage is exactly what recursion is bad at: the agent reads region A, then region B that overlaps A, then part of A again, and every level re-sends prior findings as prompt context. The recursive run cost 41,500 tokens and 34.6 s against the baseline's 34,800 and 8.2 s — the same accuracy for roughly 4x the latency. If your question is "what does this file do overall," embed it and take the top-k.

The Silent Context Budget Drain

This is the failure mode that makes recursive retrieval look cheap in a demo and expensive in production. Transcript growth is superlinear in depth, because every level re-includes prior findings as prompt context. Per-call curve for the synthesis question that hit the budget stop:

DepthPrompt tokens billedCumulative billed
0612612
11,3882,000
23,9045,904
39,18015,084
421,46036,544

That's roughly 2.3x per level, and it's why depth limits that feel generous in a prototype (10, 15) are unusable on a 45,000-token budget. The budget check rejected the depth-5 call; without it, the run would have continued to a depth-5 prompt around 50,000 tokens on its own.

The second drain is unbounded tool output, and the minified bundle makes it concrete. A grep for retryBudget against vendor.bundle.js returns one line of 480,113 bytes. Uncapped, that's roughly 120,000 tokens — an estimate from a byte-count division, but the order of magnitude is the point. With the 8,192-byte cap at the tool boundary, the same call returns about 2,000 tokens and a truncated: true flag. Capping inside the tool, not in the prompt template, is the defense: the model never sees the full line, so it can't accidentally forward it.

Four defenses, all measured as working in the harness:

  • Dedupe findings by path:start-end so the agent can't re-read the same region under a slightly different question.
  • Summarize between levels instead of forwarding raw tool output; compact(question, findings) replaces the transcript each turn.
  • Cap tool output at the tool boundary with both a line cap and a byte cap, and return a truncation flag so the model knows to narrow.
  • Log token spend per level, because the curve above is invisible if you only record a run total.

Reproducing the Benchmark Harness

node harness/run.mjs --pipeline=baseline  --queries=fixtures/queries.jsonl \
  --chunk=1200 --overlap=200 --k=8 --out=results/baseline.jsonl

node harness/run.mjs --pipeline=recursive --queries=fixtures/queries.jsonl \
  --max-depth=6 --max-tokens=45000 --out=results/recursive.jsonl

node harness/report.mjs results/*.jsonl

report.mjs prints one row per pipeline/query-type cell with the same columns as the table above, plus the per-depth token curve for any run that stopped on budget. Sample output from the symbol cells:

pipeline   queryType  n  tokens  meanMs  correct
baseline   symbol     3   28400    6100    2/3
recursive  symbol     3    3900    4800    3/3
💪

Pin the chat model version and the embedding model version in config before you compare anything. Neither pipeline's numbers transfer across model versions, and an unpinned comparison measures the provider's release schedule, not your retrieval design.

A Decision Rule for Choosing Between Recursive and Vector Search

Query shapeCorpus shapeUse
Exact symbol or lineAny sizeRecursive grep + narrow read
Cross-region trace / "where is X read and written"Single huge fileRecursive, with dedupe by line range
Broad synthesisOne file, fits in a few k chunksBaseline chunk-and-embed
Broad synthesisWhole repoNeither alone — vector prefilter, then recursive read
Repeated identical questionsStable corpusBaseline; cache the index and the answers
Tight token budget, high query volumeEitherBaseline, then escalate to recursive on low-confidence answers

The "neither, use both" row is the one most teams land on and the one I did not implement here. Vector top-k as a prefilter to pick the file or region, then recursive reads inside it, is the shape that scales to a real repo. That row is untested inference on my part, and the harness has no hybrid pipeline to back it.

Limitations

Single machine, one chat model version, one embedding model version, one run per query, nine queries total. I wrote the fixtures, and I wrote the answer key — so I could have unconsciously written symbols that favour grep. The synthesis rubric is binary and I scored it by reading the transcripts, not by a second reviewer.

What's measured: the token ledger, the latency, the correctness column above, the per-depth token curve, and the 480 KB single-line grep result. What's extrapolation: the decision table, the "hybrid is better at repo scale" claim, and any statement about how these numbers would move with a different model version. Reproduce those before you plan capacity around them.

Further Reading

Share this post

More posts

Comments