
Metering Always-On Agents in Node.js: Cost per Task Across GPT-6.1 Sol and Astra
Introduction: Why Always-On Agents Need Cost-per-Task Metering
Always-on agents break the request/response cost model. A cloud-hosted agent that runs continuously doesn't bill per feature — it bills for every heartbeat, every retry, and every turn of a conversation nobody is reading. This post shows how to meter a Node.js agent per task so you can compare real cost per task across GPT-6.1 Sol, GPT-6 Astra, and the 300 tok/s Ultrafast tier instead of budgeting on a launch headline. OpenAI's DevDay 2026 coverage is a fine excuse for the exercise, but the coverage isn't the point — the accounting is.
Below: what the published reporting actually says and what it leaves out, a metering contract for agent tasks, a provider-agnostic wrapper in Node.js, a worked cost-per-task calculation with a break-even rule, and why the 300 tok/s Ultrafast tier is a latency claim rather than a cost claim.
What the DevDay 2026 Coverage Actually Claims About GPT-6.1 Sol and Astra — and What It Doesn't
Confirmed from the published reporting
- On 2026-09-29, marktechpost reported that OpenAI launched dots, described as always-on GPT-6 Astra agents working from their own cloud computers.
- The same day, VentureBeat reported that GPT-6.1 Sol offers Astra-like performance at roughly one-fifth the price, alongside a new Ultrafast tier at 300 tokens per second.
- suaragarut.id reported the same launch at DevDay 2026.
- OpenTools reported a Codex CLI refresh for voice and agent oversight, and noted that the underlying features predate DevDay.
"Confirmed" here means an outlet published it. It does not mean OpenAI's benchmark reproduces.
Unverified or missing from the public snippets
- No numeric rate card. "One-fifth the price" doesn't say which token classes are discounted — input, cached input, output, reasoning — or by how much. Untested.
- No evaluation methodology behind "Astra-like performance": the snippets carry no task set, no sample size, no pass criterion.
- No token-comparability claim. A model that is 5x cheaper per token but needs 1.6x the tokens is about 3.1x cheaper per task.
- No product linkage details: whether dots runs on Sol, Astra, or both; concurrency limits; whether Ultrafast carries its own rate.
- No primary announcement or full pricing page in the material I had — only secondary coverage.
That gap is why the rest of this post is about measurement rather than a verdict.
Why Cost per Task Beats Cost per Token for Always-On Agents
The cost-per-success formula that matters
cost_per_success = (sum of all attempt costs + infra share + tool share) / tasks_succeeded
Per-token price is an input to that formula, not an output. Finance reads the left side.
Why idle and retry traffic breaks per-request intuition
An always-on agent emits sessions that never become tasks: heartbeats, polling, "check if anything changed." Those cost money with a denominator of zero. Retries cost twice — the failed attempt and the successful one — and both look identical in a per-request dashboard. Cost per task is the only metric that forces you to count them.
The Metering Contract: What to Capture for Every Task
| Field group | Fields | Why it exists |
|---|---|---|
| Tokens | input_tokens, output_tokens, cached_input_tokens, reasoning_tokens | Rate cards price these separately |
| Identity | task_id, parent_task_id, agent_id, tenant, task_class, trace_id | Attribution and later routing |
| Time | wall_ms, model_ms, tool_ms, queue_ms | Separates model latency from everything else |
| Outcome | outcome, error_class, attempts, escalated_from | Joins cost to success |
| Rate card | rate_card_version | Makes old records reproducible |
Token accounting
cached_input_tokens is normally a subset of input_tokens, not an additional count. Subtract before billing the uncached rate, or you double-bill your own metrics. Reasoning tokens are reported separately but billed as output on OpenAI's reasoning models per the API docs; confirm the same for Sol and Ultrafast rather than assuming it.
Task identity and attribution keys
Without tenant and task_class you can't compute cost per customer, and you can't route by workload type later. Add them on day one — backfilling is painful.
Wall clock, tool latency, and model latency split
A faster tier only moves model_ms. If tool_ms dominates your p95, paying for speed is paying for the wrong leg.
Outcome flags so cost can be joined to success
Use a closed set: success, task_failure, error, abandoned. Free-text outcome strings make aggregation impossible within a week.
Building a Provider-Agnostic Metering Wrapper in Node.js
A MeteredAgent wrapper around the model client
import { appendFile } from "node:fs/promises";
// Placeholder rate card: USD per 1M tokens.
// Replace with your provider's published prices and version this file.
export const RATE_CARD_VERSION = "2026-10-01.placeholder";
const RATE_CARD = {
"astra-class": { input: 15, cachedInput: 1.5, output: 75 },
"sol-class": { input: 3, cachedInput: 0.3, output: 15 },
ultrafast: { input: 3, cachedInput: 0.3, output: 15 },
};
export function costUsd(tier, usage) {
const r = RATE_CARD[tier];
const cached = usage.cached_input_tokens ?? 0;
const uncached = Math.max(usage.input_tokens - cached, 0);
return (
(uncached * r.input) / 1e6 +
(cached * r.cachedInput) / 1e6 +
((usage.output_tokens ?? 0) * r.output) / 1e6
);
}
export function createMeteredAgent({ client, tier, sinkPath, now = Date.now }) {
return async function runTask(task, attempt) {
const t0 = now();
const record = {
task_id: task.id,
tenant: task.tenant ?? "internal",
task_class: task.class,
agent_id: task.agentId,
model: client.model,
tier,
rate_card: RATE_CARD_VERSION,
attempts: 0,
usage: { input_tokens: 0, output_tokens: 0, cached_input_tokens: 0 },
model_ms: 0,
tool_ms: 0,
wall_ms: 0,
outcome: "error",
cost_usd: 0,
};
try {
const result = await attempt({
onAttempt: () => { record.attempts += 1; },
onUsage: (u) => {
record.usage.input_tokens += u.input_tokens ?? 0;
record.usage.output_tokens += u.output_tokens ?? 0;
record.usage.cached_input_tokens += u.cached_input_tokens ?? 0;
},
onLatency: (l) => {
record.model_ms += l.model ?? 0;
record.tool_ms += l.tool ?? 0;
},
});
record.outcome = result.ok ? "success" : "task_failure";
record.error_class = result.errorClass;
} catch (err) {
record.outcome = "error";
record.error_class = err?.name ?? "UnknownError";
} finally {
record.wall_ms = now() - t0;
record.cost_usd = Number(costUsd(tier, record.usage).toFixed(6));
await appendFile(sinkPath, JSON.stringify(record) + "
");
}
return record;
};
}Failed attempts still get metered. If the provider returns no usage on a mid-stream failure, record usage_partial: true instead of zero — a silent zero is worse than an estimate.
Capturing usage from streamed responses without losing the final chunk
export async function collectStream(stream, bucket) {
const t0 = Date.now();
let text = "";
let usage = null;
for await (const chunk of stream) {
const delta = chunk.choices?.[0]?.delta?.content;
if (delta) text += delta;
// The usage chunk arrives AFTER the finish_reason chunk and carries an
// empty choices array. Breaking on finish_reason silently loses it.
if (chunk.usage) usage = chunk.usage;
}
bucket.model += Date.now() - t0;
return { text, usage };
}
const stream = await client.chat.completions.create({
model: client.model,
messages,
stream: true,
stream_options: { include_usage: true },
});stream_options.include_usage is documented in the OpenAI Chat Completions API reference. The agent-side bug I see most often is an early break on finish_reason, which produces a record with zero tokens and a nonzero bill.
Writing task records to an append-only JSONL sink
appendFile per task holds up to a few thousand records per second. One record per line, never rewritten, deduplicated on task_id at read time. Keep prompt text out of the sink — redact to a hash plus length — because these records double as your audit trail for what an always-on agent did.
A Worked Cost-per-Task Calculation for Sol, Astra, and Ultrafast
Price inputs and how to source them correctly
Take prices from the provider's published pricing page or your invoice, never from launch coverage. Store them as a versioned rate card committed to the repo, assert on the version at startup, and reconcile monthly totals against the invoice line. The numbers below are illustrative placeholders; the arithmetic is the transferable part.
Three model tiers compared on the same task record
Workload: a 12-turn triage task. Turn 1 is 2,200 input tokens (0 cached) and 300 output. Turns 2–12 are 2,600 input tokens each, of which 1,800 are cached, plus 250 output. Totals per task: 30,800 input (19,800 cached, 11,000 uncached) and 3,050 output.
| Tier | Uncached in | Cached in | Output | $/task |
|---|---|---|---|---|
| astra-class (15 / 1.5 / 75) | $0.165 | $0.0297 | $0.2288 | $0.4235 |
| sol-class (3 / 0.3 / 15) | $0.033 | $0.0059 | $0.0458 | $0.0847 |
| ultrafast (assumed = sol) | $0.033 | $0.0059 | $0.0458 | $0.0847 |
The ratio is exactly 5.0 — but only because I held token counts constant and applied 0.2 to every token class. Both assumptions are unverified for Sol.
Reading the break-even point instead of trusting the headline multiple
With these placeholders, Sol wins per task while it uses less than 5x Astra's tokens. If Sol burns 1.6x the tokens on every class, the multiple becomes 5 / 1.6 ≈ 3.1x. Add failures and it drops again: at 187/200 successes for Sol versus 194/200 for Astra, cost per success is $0.0906 versus $0.4366 — a 4.8x multiple, not 5x.
The Ultrafast Tier Trade-off: 300 tok/s Is a Latency Claim, Not a Cost Claim
What generation speed does and does not change about your bill
Throughput doesn't change cost per token unless Ultrafast has its own rate card, which the published snippets don't state. It shortens the completion leg of model_ms and nothing else in the formula.
Where a faster tier actually pays for itself in an agent loop
- Human-in-the-loop turns, where wall clock is the product.
- Timeout-driven retries. If a turn runs near your client timeout, a 3x faster completion converts retries into successes — a real token saving, not a rounding error.
- Concurrency: faster completions mean fewer simultaneous workers and less idle container time on the agent host.
Buy speed for interactive classes and timeout-prone ones. Don't buy it for nightly batch work.
Measuring It in Your Own Workload: A Small Honest Harness
Reproducible run command and sample size
node bench/run-agent-cost.mjs \
--tasks bench/tasks/triage-200.jsonl \
--tiers astra-class,sol-class,ultrafast \
--provider mock-replay \
--out runs/2026-10-01
This run replays a stored usage trace, so the arithmetic is deterministic. I have not run it against real Sol or Ultrafast endpoints; treat the model rows as simulated. 200 tasks, one repeat — enough to catch anything wrong by a factor of two, not enough to resolve a 5% difference.
tier attempts success failed tokens_in tokens_out cached cost_usd usd_per_success
sol-class 200 187 13 6.16M 0.61M 64% 16.94 0.0906
astra-class 200 194 6 6.16M 0.61M 64% 84.69 0.4366
ultrafast 200 190 10 6.16M 0.61M 64% 16.94 0.0892
Reporting results with counts, not percentages alone
"187/200 succeeded; 13 failed; 9 of the 13 were tool-schema validation errors" beats "93.5% success rate." The denominator and the failure class are the actionable parts. Always report the tier, the criterion, and the rate-card version next to the number.
Where Always-On Agent Cost Actually Leaks
Context growth across turns
A naive agent resends the full transcript, so turn 12 costs roughly 12x turn 1 in input tokens, and every turn pays for the system prompt again. Cache the stable prefix and trim or summarize older turns. Watch cached_input_tokens / input_tokens; under 50% on a long-running agent means that's your first fix.
Retry storms and duplicate tool calls
Retries without idempotency keys execute the same write twice and pay for two model attempts. Track attempts per task and alarm on p95.
Heartbeat work that produces no task
| Heartbeat design | Tokens/day | $/day (Sol placeholders) | $/month |
|---|---|---|---|
| Model call every 30s (1,200 in / 40 out) | 3.46M in, 0.115M out | $12.10 | ~$363 |
| Same, prompt cached | 3.46M cached, 0.115M out | $2.77 | ~$83 |
| Deterministic check, model only on change (10 calls/day) | 12k in, 400 out | $0.04 | ~$1.26 |
That first row is one agent doing nothing for the price of a small instance. Caching the heartbeat prompt cuts it by 77%; replacing the heartbeat with an event cuts it by ~99%. Do the second one.
Routing Policy: When the Cheap Tier Is Enough and When the Expensive One Pays
Escalation rules based on task class, not vibes
| Task class | Default tier | Escalation trigger |
|---|---|---|
| Classify, route, extract against a known schema | sol-class | Schema validation fails twice |
| Code edit with a test suite | sol-class | Tests fail after the generated patch |
| Ambiguous planning, multi-repo changes | astra-class | — |
| Long-context review with no validator | astra-class | — |
Trigger on evidence a validator produced, never on the agent's self-reported confidence.
Guardrails that keep a router from silently upgrading everything
- One escalation per task, recorded as
escalated_fromwithescalation_reason. - Per-task token cap; abort to
outcome: "abandoned"when exceeded. - Alert when the escalation share exceeds 20% week over week.
- Shadow-run the cheap tier on 5% of tasks routed to the expensive tier and compare outcomes — otherwise your router is never falsified.
My Position: Instrument First, Route Second, and Do Not Budget on a Press Multiple
The 5x figure is a per-token price claim from secondary coverage, not a cost-per-task result, and the gap between 5x and the ~3.1x you get from a 1.6x token blowup is a real budget line. Build the task record before you build the router: it's a day of work, and it's the only instrument that tells you which multiple you're actually getting. Routing without metering is guessing with extra steps.
One more thing worth saying plainly: always-on agents running on their own cloud computers widen the blast radius, and the same JSONL that produces your invoice is the audit trail for what those agents did. Keep tool calls, outcomes, and timings. Redact prompt content by default.
Further Reading
- OpenAI API pricing — the rate card to source token prices from directly.
- OpenAI Chat Completions API reference —
stream_options.include_usageand the usage object. - VentureBeat coverage of GPT-6.1 Sol and the Ultrafast tier, 2026-09-29 — source of the one-fifth-price and 300 tok/s claims (Google News feed link as provided).
- marktechpost — reported the 2026-09-29 launch of dots as always-on GPT-6 Astra agents on their own cloud computers (secondary coverage; outlet homepage, since the original feed link was not preserved).
Share this post
More posts

Measuring Tokens-per-Task in JavaScript Agents with Per-Step Cost Tracing and Budget Limits

How to Harden Your AI Agent After the Mythos and Fable Prompt Injection Breach
