Metering Always-On Agents in Node.js: Cost per Task Across GPT-6.1 Sol and Astra

Metering Always-On Agents in Node.js: Cost per Task Across GPT-6.1 Sol and Astra

pr0h0•
nodejsai-agentsllm-cost-optimizationobservabilityopenai
AI Usage (91%)

Introduction: Why Always-On Agents Need Cost-per-Task Metering

Always-on agents break the request/response cost model. A cloud-hosted agent that runs continuously doesn't bill per feature — it bills for every heartbeat, every retry, and every turn of a conversation nobody is reading. This post shows how to meter a Node.js agent per task so you can compare real cost per task across GPT-6.1 Sol, GPT-6 Astra, and the 300 tok/s Ultrafast tier instead of budgeting on a launch headline. OpenAI's DevDay 2026 coverage is a fine excuse for the exercise, but the coverage isn't the point — the accounting is.

Below: what the published reporting actually says and what it leaves out, a metering contract for agent tasks, a provider-agnostic wrapper in Node.js, a worked cost-per-task calculation with a break-even rule, and why the 300 tok/s Ultrafast tier is a latency claim rather than a cost claim.

What the DevDay 2026 Coverage Actually Claims About GPT-6.1 Sol and Astra — and What It Doesn't

Confirmed from the published reporting

  • On 2026-09-29, marktechpost reported that OpenAI launched dots, described as always-on GPT-6 Astra agents working from their own cloud computers.
  • The same day, VentureBeat reported that GPT-6.1 Sol offers Astra-like performance at roughly one-fifth the price, alongside a new Ultrafast tier at 300 tokens per second.
  • suaragarut.id reported the same launch at DevDay 2026.
  • OpenTools reported a Codex CLI refresh for voice and agent oversight, and noted that the underlying features predate DevDay.

"Confirmed" here means an outlet published it. It does not mean OpenAI's benchmark reproduces.

Unverified or missing from the public snippets

  • No numeric rate card. "One-fifth the price" doesn't say which token classes are discounted — input, cached input, output, reasoning — or by how much. Untested.
  • No evaluation methodology behind "Astra-like performance": the snippets carry no task set, no sample size, no pass criterion.
  • No token-comparability claim. A model that is 5x cheaper per token but needs 1.6x the tokens is about 3.1x cheaper per task.
  • No product linkage details: whether dots runs on Sol, Astra, or both; concurrency limits; whether Ultrafast carries its own rate.
  • No primary announcement or full pricing page in the material I had — only secondary coverage.

That gap is why the rest of this post is about measurement rather than a verdict.

Why Cost per Task Beats Cost per Token for Always-On Agents

The cost-per-success formula that matters

cost_per_success = (sum of all attempt costs + infra share + tool share) / tasks_succeeded

Per-token price is an input to that formula, not an output. Finance reads the left side.

Why idle and retry traffic breaks per-request intuition

An always-on agent emits sessions that never become tasks: heartbeats, polling, "check if anything changed." Those cost money with a denominator of zero. Retries cost twice — the failed attempt and the successful one — and both look identical in a per-request dashboard. Cost per task is the only metric that forces you to count them.

The Metering Contract: What to Capture for Every Task

Field groupFieldsWhy it exists
Tokensinput_tokens, output_tokens, cached_input_tokens, reasoning_tokensRate cards price these separately
Identitytask_id, parent_task_id, agent_id, tenant, task_class, trace_idAttribution and later routing
Timewall_ms, model_ms, tool_ms, queue_msSeparates model latency from everything else
Outcomeoutcome, error_class, attempts, escalated_fromJoins cost to success
Rate cardrate_card_versionMakes old records reproducible

Token accounting

cached_input_tokens is normally a subset of input_tokens, not an additional count. Subtract before billing the uncached rate, or you double-bill your own metrics. Reasoning tokens are reported separately but billed as output on OpenAI's reasoning models per the API docs; confirm the same for Sol and Ultrafast rather than assuming it.

Task identity and attribution keys

Without tenant and task_class you can't compute cost per customer, and you can't route by workload type later. Add them on day one — backfilling is painful.

Wall clock, tool latency, and model latency split

A faster tier only moves model_ms. If tool_ms dominates your p95, paying for speed is paying for the wrong leg.

Outcome flags so cost can be joined to success

Use a closed set: success, task_failure, error, abandoned. Free-text outcome strings make aggregation impossible within a week.

Building a Provider-Agnostic Metering Wrapper in Node.js

A MeteredAgent wrapper around the model client

meter/metered-agent.js
import { appendFile } from "node:fs/promises";

// Placeholder rate card: USD per 1M tokens.
// Replace with your provider's published prices and version this file.
export const RATE_CARD_VERSION = "2026-10-01.placeholder";

const RATE_CARD = {
"astra-class": { input: 15, cachedInput: 1.5, output: 75 },
"sol-class": { input: 3, cachedInput: 0.3, output: 15 },
ultrafast: { input: 3, cachedInput: 0.3, output: 15 },
};

export function costUsd(tier, usage) {
const r = RATE_CARD[tier];
const cached = usage.cached_input_tokens ?? 0;
const uncached = Math.max(usage.input_tokens - cached, 0);
return (
  (uncached * r.input) / 1e6 +
  (cached * r.cachedInput) / 1e6 +
  ((usage.output_tokens ?? 0) * r.output) / 1e6
);
}

export function createMeteredAgent({ client, tier, sinkPath, now = Date.now }) {
return async function runTask(task, attempt) {
  const t0 = now();
  const record = {
    task_id: task.id,
    tenant: task.tenant ?? "internal",
    task_class: task.class,
    agent_id: task.agentId,
    model: client.model,
    tier,
    rate_card: RATE_CARD_VERSION,
    attempts: 0,
    usage: { input_tokens: 0, output_tokens: 0, cached_input_tokens: 0 },
    model_ms: 0,
    tool_ms: 0,
    wall_ms: 0,
    outcome: "error",
    cost_usd: 0,
  };

  try {
    const result = await attempt({
      onAttempt: () => { record.attempts += 1; },
      onUsage: (u) => {
        record.usage.input_tokens += u.input_tokens ?? 0;
        record.usage.output_tokens += u.output_tokens ?? 0;
        record.usage.cached_input_tokens += u.cached_input_tokens ?? 0;
      },
      onLatency: (l) => {
        record.model_ms += l.model ?? 0;
        record.tool_ms += l.tool ?? 0;
      },
    });
    record.outcome = result.ok ? "success" : "task_failure";
    record.error_class = result.errorClass;
  } catch (err) {
    record.outcome = "error";
    record.error_class = err?.name ?? "UnknownError";
  } finally {
    record.wall_ms = now() - t0;
    record.cost_usd = Number(costUsd(tier, record.usage).toFixed(6));
    await appendFile(sinkPath, JSON.stringify(record) + "
");
  }

  return record;
};
}

Failed attempts still get metered. If the provider returns no usage on a mid-stream failure, record usage_partial: true instead of zero — a silent zero is worse than an estimate.

Capturing usage from streamed responses without losing the final chunk

meter/stream-usage.js
export async function collectStream(stream, bucket) {
const t0 = Date.now();
let text = "";
let usage = null;

for await (const chunk of stream) {
  const delta = chunk.choices?.[0]?.delta?.content;
  if (delta) text += delta;

  // The usage chunk arrives AFTER the finish_reason chunk and carries an
  // empty choices array. Breaking on finish_reason silently loses it.
  if (chunk.usage) usage = chunk.usage;
}

bucket.model += Date.now() - t0;
return { text, usage };
}

const stream = await client.chat.completions.create({
model: client.model,
messages,
stream: true,
stream_options: { include_usage: true },
});

stream_options.include_usage is documented in the OpenAI Chat Completions API reference. The agent-side bug I see most often is an early break on finish_reason, which produces a record with zero tokens and a nonzero bill.

Writing task records to an append-only JSONL sink

appendFile per task holds up to a few thousand records per second. One record per line, never rewritten, deduplicated on task_id at read time. Keep prompt text out of the sink — redact to a hash plus length — because these records double as your audit trail for what an always-on agent did.

A Worked Cost-per-Task Calculation for Sol, Astra, and Ultrafast

Price inputs and how to source them correctly

Take prices from the provider's published pricing page or your invoice, never from launch coverage. Store them as a versioned rate card committed to the repo, assert on the version at startup, and reconcile monthly totals against the invoice line. The numbers below are illustrative placeholders; the arithmetic is the transferable part.

Three model tiers compared on the same task record

Workload: a 12-turn triage task. Turn 1 is 2,200 input tokens (0 cached) and 300 output. Turns 2–12 are 2,600 input tokens each, of which 1,800 are cached, plus 250 output. Totals per task: 30,800 input (19,800 cached, 11,000 uncached) and 3,050 output.

TierUncached inCached inOutput$/task
astra-class (15 / 1.5 / 75)$0.165$0.0297$0.2288$0.4235
sol-class (3 / 0.3 / 15)$0.033$0.0059$0.0458$0.0847
ultrafast (assumed = sol)$0.033$0.0059$0.0458$0.0847

The ratio is exactly 5.0 — but only because I held token counts constant and applied 0.2 to every token class. Both assumptions are unverified for Sol.

Reading the break-even point instead of trusting the headline multiple

With these placeholders, Sol wins per task while it uses less than 5x Astra's tokens. If Sol burns 1.6x the tokens on every class, the multiple becomes 5 / 1.6 ≈ 3.1x. Add failures and it drops again: at 187/200 successes for Sol versus 194/200 for Astra, cost per success is $0.0906 versus $0.4366 — a 4.8x multiple, not 5x.

The Ultrafast Tier Trade-off: 300 tok/s Is a Latency Claim, Not a Cost Claim

What generation speed does and does not change about your bill

Throughput doesn't change cost per token unless Ultrafast has its own rate card, which the published snippets don't state. It shortens the completion leg of model_ms and nothing else in the formula.

Where a faster tier actually pays for itself in an agent loop

  • Human-in-the-loop turns, where wall clock is the product.
  • Timeout-driven retries. If a turn runs near your client timeout, a 3x faster completion converts retries into successes — a real token saving, not a rounding error.
  • Concurrency: faster completions mean fewer simultaneous workers and less idle container time on the agent host.

Buy speed for interactive classes and timeout-prone ones. Don't buy it for nightly batch work.

Measuring It in Your Own Workload: A Small Honest Harness

Reproducible run command and sample size

node bench/run-agent-cost.mjs \
  --tasks bench/tasks/triage-200.jsonl \
  --tiers astra-class,sol-class,ultrafast \
  --provider mock-replay \
  --out runs/2026-10-01

This run replays a stored usage trace, so the arithmetic is deterministic. I have not run it against real Sol or Ultrafast endpoints; treat the model rows as simulated. 200 tasks, one repeat — enough to catch anything wrong by a factor of two, not enough to resolve a 5% difference.

tier         attempts  success  failed  tokens_in  tokens_out  cached  cost_usd  usd_per_success
sol-class    200       187      13      6.16M      0.61M       64%     16.94     0.0906
astra-class  200       194      6       6.16M      0.61M       64%     84.69     0.4366
ultrafast    200       190      10      6.16M      0.61M       64%     16.94     0.0892

Reporting results with counts, not percentages alone

"187/200 succeeded; 13 failed; 9 of the 13 were tool-schema validation errors" beats "93.5% success rate." The denominator and the failure class are the actionable parts. Always report the tier, the criterion, and the rate-card version next to the number.

Where Always-On Agent Cost Actually Leaks

Context growth across turns

A naive agent resends the full transcript, so turn 12 costs roughly 12x turn 1 in input tokens, and every turn pays for the system prompt again. Cache the stable prefix and trim or summarize older turns. Watch cached_input_tokens / input_tokens; under 50% on a long-running agent means that's your first fix.

Retry storms and duplicate tool calls

Retries without idempotency keys execute the same write twice and pay for two model attempts. Track attempts per task and alarm on p95.

Heartbeat work that produces no task

Heartbeat designTokens/day$/day (Sol placeholders)$/month
Model call every 30s (1,200 in / 40 out)3.46M in, 0.115M out$12.10~$363
Same, prompt cached3.46M cached, 0.115M out$2.77~$83
Deterministic check, model only on change (10 calls/day)12k in, 400 out$0.04~$1.26

That first row is one agent doing nothing for the price of a small instance. Caching the heartbeat prompt cuts it by 77%; replacing the heartbeat with an event cuts it by ~99%. Do the second one.

Routing Policy: When the Cheap Tier Is Enough and When the Expensive One Pays

Escalation rules based on task class, not vibes

Task classDefault tierEscalation trigger
Classify, route, extract against a known schemasol-classSchema validation fails twice
Code edit with a test suitesol-classTests fail after the generated patch
Ambiguous planning, multi-repo changesastra-class—
Long-context review with no validatorastra-class—

Trigger on evidence a validator produced, never on the agent's self-reported confidence.

Guardrails that keep a router from silently upgrading everything

  • One escalation per task, recorded as escalated_from with escalation_reason.
  • Per-task token cap; abort to outcome: "abandoned" when exceeded.
  • Alert when the escalation share exceeds 20% week over week.
  • Shadow-run the cheap tier on 5% of tasks routed to the expensive tier and compare outcomes — otherwise your router is never falsified.

My Position: Instrument First, Route Second, and Do Not Budget on a Press Multiple

The 5x figure is a per-token price claim from secondary coverage, not a cost-per-task result, and the gap between 5x and the ~3.1x you get from a 1.6x token blowup is a real budget line. Build the task record before you build the router: it's a day of work, and it's the only instrument that tells you which multiple you're actually getting. Routing without metering is guessing with extra steps.

One more thing worth saying plainly: always-on agents running on their own cloud computers widen the blast radius, and the same JSONL that produces your invoice is the audit trail for what those agents did. Keep tool calls, outcomes, and timings. Redact prompt content by default.

Further Reading

Share this post

More posts

Comments