Rebuilding Routing Thresholds and RAG Fan-Out for Claude Haiku 5.5's New Pricing

Rebuilding Routing Thresholds and RAG Fan-Out for Claude Haiku 5.5's New Pricing

pr0h0•
claude-haikullm-pricingragmodel-routinganthropic
AI Usage (87%)

Why Claude Haiku 5.5's Pricing Reset Forces a Routing Rebuild

Coverage dated 2026-10-07 and 2026-10-08 puts Claude Haiku 5.5 at roughly $0.10 per million input tokens and $0.50 per million output tokens, an average cut near 75% versus its predecessor. This post works out what that cut does to the two things you actually control: the threshold at which your router escalates to a larger model, and the RAG fan-out your pipeline can afford per query. The four pieces I have — 36Kr, NeoTeo, Forkast News, and The Tech Portal — all read it as a small-model price war; 36Kr calls it the move into the "1-cent era."

My position: a small-tier price cut is a routing-config change, not a re-benchmark story. It doesn't change what your models should do. It changes where the line between them sits, and it moves that line by a measurable factor. Any cheap-model threshold you hardcoded last quarter is now wrong in a specific, arithmetic way — the small tier can absorb about four times more work per dollar, so every threshold copied from an old spreadsheet now under-serves it.

One caveat before the numbers. What I have is aggregated headline snippets from Google News, not Anthropic's pricing page. That limits what I can assert, and below I've kept the confirmed parts visibly separate from the inferred ones.

What the Pricing Coverage Confirms and What It Leaves Unknown

What the cited sources report on Haiku 5.5 pricing

  • 36Kr (2026-10-08, 01:41 UTC): Haiku 5.5 "cuts prices to rock-bottom levels," framed as small-model competition entering the 1-cent era.
  • NeoTeo (2026-10-08, 00:43 UTC): headline is "Claude Haiku 5.5 pricing: API rates by prompt size" — the only signal here that rates vary by prompt size.
  • Forkast News (2026-10-07, 23:53 UTC): "$0.10/$0.50" stated directly, described as collapsing the small-model pricing floor, "eight days before the October 15 deadline."
  • The Tech Portal (2026-10-07, 20:48 UTC): Haiku 5.5 as Anthropic's most capable small model yet and far cheaper than its predecessor.

Treat as snippet-confirmed: the $0.10/$0.50 pair, an average drop near 75%, prompt-size tiering, a launch window around 2026-10-07/08, and the existence of an October 15 deadline. Forkast's timestamp and its "eight days" framing are internally consistent — 2026-10-07 plus eight days is 2026-10-15.

What the cited snippets do not establish

QuestionStatus from these sources
The exact predecessor rate the 75% is measured againstunknown
Prompt-size tier boundaries, and the rate above themunknown
Whether batch or cache discounts stack on these ratesunknown
Whether older Haiku SKUs are deprecatedunknown
What the October 15 deadline actually isunknown

The predecessor rate deserves one paragraph of arithmetic, because the headline numbers don't obviously reconcile. A uniform 75% cut from $0.40/$2.00 lands exactly on $0.10/$0.50. But 75% off $1/$5 would land on $0.25/$1.25, and a cut to $0.10/$0.50 from $1/$5 is 90%, not 75%. So one of three things holds: the predecessor was around $0.40/$2.00, the reported rates are rounded, or "average near 75%" is weighted by a token mix rather than a plain mean of the two per-token cuts. Headlines don't tell me which. It matters, because it is the difference between a 4x budget increase and a 10x one.

💡

Verify rates, tier boundaries, and any batch or cache multipliers on Anthropic's official pricing page before you ship a threshold change. The numbers in this post are the reported ones from the cited coverage.

The Worked Crossover: Why Your Routing Threshold Just Moved

Price cost per request, not cost per token

Per-token prices are a distraction. Your router decides whether this request is worth escalating. So price one request end to end. Assumption, stated so you can substitute yours: a 400-token router/classifier prompt with 30 output tokens, plus a 2,400-token task prompt carrying retrieved context, with 350 output tokens. Total 2,800 in / 380 out, identical for both tiers.

TierRate (per 1M in/out)Input costOutput costCost / requestPer 1,000 requests
Small, inferred predecessor$0.40 / $2.00$0.00112$0.00076$0.00188$1.88
Small, reported Haiku 5.5$0.10 / $0.50$0.00028$0.00019$0.00047$0.47
Large fallback (your number)$3.00 / $15.00$0.00840$0.00570$0.01410$14.10

The last column is the one that matters. A request that cost 7.5x the cheap tier now costs 30x it. At $0.00047, one routed request is under half a cent — on this prompt shape, 36Kr's "1-cent era" framing is arithmetically fair.

What the pricing shift actually changes

Nothing about the model. Everything about the budget. If the small tier is ~4x cheaper per call, the same dollar buys roughly 4x more of anything that scales per call: retries with a repair prompt, few-shot examples in the classifier, a larger retrieved context, or a second pass to check the first. All of those are knobs to move before you consider pushing requests up to the bigger model.

The concrete move: if your rule was "escalate when the small model's confidence is below 0.7," run the retry ladder instead. Four cheap attempts at $0.00047 cost $0.00188 — exactly one old-priced call, and still 7.5x cheaper than a single large-model attempt.

Measure your own traffic instead of re-guessing

Derive the threshold from your own traffic. Log tokens, then recompute.

route-cost.mjs
// node route-cost.mjs logs/*.jsonl
// Expects one JSON object per call:
// { decision, tier, model, latency_ms, usage: { input_tokens, output_tokens,
//   cache_read_input_tokens, cache_creation_input_tokens } }


// usd per 1M tokens. Fill from the official pricing page.
// Cache multipliers here are placeholders, not confirmed for Haiku 5.5.
const RATE = {
small: { input: 0.10, output: 0.50, cacheRead: 0.010, cacheWrite: 0.125 },
large: { input: 3.00, output: 15.00, cacheRead: 0.300, cacheWrite: 3.750 },
};

function costPerCall(u, r) {
const {
  input_tokens = 0,
  output_tokens = 0,
  cache_read_input_tokens = 0,
  cache_creation_input_tokens = 0,
} = u ?? {};
return (
  input_tokens * r.input +
  output_tokens * r.output +
  cache_read_input_tokens * r.cacheRead +
  cache_creation_input_tokens * r.cacheWrite
) / 1e6;
}

const rows = process.argv.slice(2)
.flatMap((f) => readFileSync(f, "utf8").trim().split("
"))
.filter(Boolean)
.map((l) => JSON.parse(l));

const agg = new Map();
for (const row of rows) {
const key = row.decision + "|" + row.model;
const b = agg.get(key) ?? { calls: 0, usd: 0, inTok: 0, outTok: 0, ms: [] };
b.calls += 1;
b.usd += costPerCall(row.usage, RATE[row.tier]);
b.inTok += row.usage?.input_tokens ?? 0;
b.outTok += row.usage?.output_tokens ?? 0;
b.ms.push(row.latency_ms);
agg.set(key, b);
}

console.table(
[...agg].map(([key, b]) => {
  const sorted = b.ms.sort((x, y) => x - y);
  return {
    key,
    calls: b.calls,
    usdPerCall: Number((b.usd / b.calls).toFixed(6)),
    avgInTok: Math.round(b.inTok / b.calls),
    avgOutTok: Math.round(b.outTok / b.calls),
    p95ms: sorted[Math.floor(sorted.length * 0.95)],
  };
})
);

The output is a table you can read directly: usdPerCall per decision, plus p95 latency. If usdPerCall for the escalated bucket is more than ~20x the non-escalated bucket, your threshold is too conservative at current prices.

RAG Fan-Out: Where the Cheap Tier Actually Pays Off

Fan-out multiplies calls, so per-call price dominates total cost far more aggressively than it does in a single-shot pipeline. That is where the reported cut changes architecture rather than arithmetic.

The RAG fan-out cost model

One query decomposes into N query variants, M retrieved chunks, one rerank or score call per chunk, and one synthesis call. The rewrite stage scales with N. The rerank stage scales with N × M. Synthesis is one call regardless. So the two stages that benefit from a 4x cheaper tier are exactly the two that used to be too expensive to run at scale.

⚠️

A headline rate cut is not a quality guarantee. Re-run your own eval set against Haiku 5.5 before lowering a routing threshold — a cheaper misroute is still a misroute.

Retrieve-then-rerank: widening the candidate set

With a cheap scorer, widening M stops being a budget question. Reranking 40 chunks used to cost real money; now it costs fractions of a cent, and widening to 100 chunks is a linear move, not a step change. State the tradeoff honestly: the win is dollars; the losses are latency and rate limits. Four times the scoring calls is four times the wall-clock fan-out and four times the request volume against a provider limit you do not control. And a cost win only counts if answer quality holds — reranking with a smaller model can reshuffle your top-k in ways that hurt synthesis, which is why the replay harness below is not optional.

RAG fan-out budget table: where the money sits

Same query shape, N = 4, M = 40, priced at the reported small rate. Synthesis is shown at both tiers so you can see where the money sits.

StageCalls / queryTokens / call (in / out)Cost @ $0.10/$0.50Cost @ $3.00/$15.00
Query rewrite4600 / 80$0.00040$0.01200
Rerank / score401,200 / 20$0.00520$0.15600
Synthesis18,000 / 700$0.00115$0.03450
Total45—$0.00675—

Two readings. First, demoting rewrite and rerank to the small tier costs $0.0056 per query — under six-tenths of a cent — and that number was four times larger before the cut. Second, and more important: if you leave synthesis on the large model, synthesis alone is $0.0345, roughly 5x the entire small-model fan-out combined. The cut doesn't change which stage deserves the expensive model. It changes how much fan-out you can afford around it.

Agent Loops and Batch Jobs: Compounding Per-Call Cost

Loop iteration cost ceiling on a fixed budget

Agent loops call a small model thousands of times per task, so per-call price compounds. Assume 40 iterations per task, each with one planning call, one tool-result summarization call, and 1.3 retries average — call it 3.3 calls per iteration, 132 calls per task, at 2,000 input / 150 output tokens each.

Rate assumptionCost per taskCalls affordable on a $0.15 budget
Inferred predecessor ($0.40/$2.00)$0.1452~132
Reported Haiku 5.5 ($0.10/$0.50)$0.0363~530

The ceiling is the actionable number. If your loop was capped at 40 iterations because 60 blew the budget, the same budget now covers roughly 530 calls — about 160 iterations at the same call density.

Batch discounts and prompt caching caveats

These three multiply, so don't average them:

  • Prompt-size tiers. NeoTeo's headline is specifically about rates by prompt size. If that's real, a long-context agent loop doesn't sit in the $0.10/$0.50 band at all, and a blended per-token average understates your bill on exactly the workloads this post is about.
  • Batch discounts. Apply only if your workload is genuinely asynchronous.
  • Prompt caching. Cache reads and writes are priced differently from uncached tokens, and the multipliers are not established for Haiku 5.5 from these sources.

Because of the tiering, segment your cost estimates by prompt-length bucket rather than averaging across them. A p50 of 1,800 tokens and a p95 of 40,000 can land in two different bands, and the p95 is where the money goes.

A Re-Evaluation Harness You Can Run Against Your Own Traffic

Instrument per-call cost and the routing decision first

Log per call: model id, input_tokens, output_tokens, cached read and write tokens, latency_ms, and — critically — the routing decision. Without the decision logged, you can't attribute a cost change to a threshold change, and you'll end up guessing again.

Replay against a golden set to measure the accuracy delta

replay-golden.mjs
// node replay-golden.mjs golden.jsonl
// golden.jsonl lines: { prompt, expected }


const RATE = {
small: { input: 0.10, output: 0.50 },   // verify before trusting
large: { input: 3.00, output: 15.00 },  // your fallback
};

const cases = readFileSync(process.argv[2], "utf8")
.trim().split("
").map(JSON.parse);

async function run(model, tier, prompt) {
const r = await callModel(model, prompt); // your existing client
const usd =
  (r.usage.input_tokens * RATE[tier].input +
   r.usage.output_tokens * RATE[tier].output) / 1e6;
return { text: r.text, usd, ms: r.latency_ms };
}

const tally = {
small: { ok: 0, usd: 0, ms: [] },
large: { ok: 0, usd: 0, ms: [] },
};

for (const c of cases) {
for (const [tier, model] of [
  ["small", "claude-haiku-5-5"],
  ["large", "your-fallback-model"],
]) {
  const r = await run(model, tier, c.prompt);
  tally[tier].usd += r.usd;
  tally[tier].ms.push(r.ms);
  if (matches(r.text, c.expected)) tally[tier].ok += 1;
}
}

for (const tier of ["small", "large"]) {
const t = tally[tier];
const s = t.ms.slice().sort((a, b) => a - b);
const p95 = s[Math.floor(s.length * 0.95)];
console.log(
  tier.padEnd(6) +
  " acc " + t.ok + "/" + cases.length +
  " | $" + t.usd.toFixed(4) + " total" +
  " | p95 " + p95 + "ms"
);
}

Run it against 50–100 stored prompts. The output format is deliberately dull:

small  acc 44/50 | $0.0231 total | p95 1180ms
large  acc 49/50 | $0.6940 total | p95 3100ms

Those numbers are a template showing the format, not measured results — I have not run Haiku 5.5, and you should not copy them. Substitute your own run. What you want to know is whether your accuracy delta is small enough to justify the price delta.

Report results next to the claims they support

Keep the result adjacent to the claim, the way you would in a test report:

Task: classify 50 support tickets into 6 routing labels.
Criterion: exact label match.
Result: small 44/50, large 49/50. All 6 small-model failures were the
same confusion pair (billing-refund vs billing-dispute).

Six failures concentrated in one confusion pair is a fixable prompt problem, not a capacity problem. Six scattered across six labels is a signal to keep the threshold where it is. And if you only have three prompts, say so — three prompts is an anecdote, not an eval.

Where Cheaper Small-Model Inference Does Not Help

Long-context and reasoning-heavy paths stay on the large tier

Heavy synthesis, code generation with large diffs, and multi-hop reasoning stay poor demotion candidates regardless of price. The failure mode is qualitative: the model doesn't produce a slightly worse answer, it produces one that is structurally wrong, and a retry at 4x the budget doesn't recover it. Cheap retries only help when the error is stochastic. Reasoning-depth errors are not.

Rate limits and latency become the bottleneck

Cheaper tokens attract traffic, and widening fan-out moves the bottleneck from budget to throughput. Going from 40 to 100 scoring calls per query doesn't just cost more — it queues more. Queue depth, p95 latency, and provider rate limits belong in the same decision as price, because a fan-out that is affordable but throttled is not affordable.

Conclusion: Recompute Routing Thresholds From Your Own Data

The reported cut to $0.10/$0.50 doesn't change what your small model should do. It changes where the routing boundary belongs, and on the prompt shape above that boundary moved roughly fourfold in the small tier's favor. Recompute it from two things: your own logged cost-per-call, and a replayed eval set that tells you the accuracy delta. Thresholds you copied from a headline — including this one — are the ones that will be wrong.

Further Reading

Share this post

More posts

Comments