← All tools

Agent loop cost calculator

An agent resends the entire conversation on every turn, so turn forty costs far more than turn one. Pricing pages quote a rate per token; this shows what that rate adds up to across a session.

Start from a workload

Large system prompt and tool definitions, file contents coming back every turn.

Per session

$2.74

Per month

$2,745

Last turn vs first

3.6×

Turn 40 costs 3.6× as much as turn 1, because every turn resends the whole conversation. Prompt caching cuts this session by 74%.

Cost of each turn — Claude Sonnet 5

Peak $0.108 per turn — hover a bar

Turn 1Turn 40

Same workload, every model

Monthly cost with caching on. Click a row to chart that model.

ModelMonthlyPer session
GPT-6 LunaOpenAI$137.24$0.137
GPT-6.1 SolOpenAI$2,298$2.30
Claude Sonnet 5Anthropic$2,745$2.74
Claude Opus 5.5Anthropic$4,596$4.60
Claude Fable 5.1Anthropic$10,373$10.37
GPT-6 AstraOpenAI$13,724$13.72
Claude Haiku 4.5AnthropicContext full at turn 35—

The point of this tool: load “Research agent”, pick GPT-6 Astra and watch the bars jump where the conversation crosses 272K tokens. Then switch caching off. The per-token price on the pricing page never changed; the bill did, because an agent is billed for its whole memory on every turn.

Why agent cost grows faster than turn count

Language model APIs are stateless: to continue a conversation you send all of it again. Turn one sends the system prompt and a message. Turn thirty sends the system prompt, twenty-nine earlier exchanges, every tool result that came back, and the new message. The input for each turn grows in a straight line, so the total for a session grows with the square of its length. Doubling the number of turns roughly quadruples the cost. Caching is what makes this affordable — the part of the conversation the model has already seen is billed at a fraction of the normal input rate — which is why the cache-read price matters more for agents than the headline input price does.

How it's calculated

  • Input on turn t = system prompt + t × tokens added per turn + (t − 1) × output per turn. Earlier outputs become part of the history.
  • With caching on, the cache-hit share of everything already seen is billed at the cache-read rate. Everything else is written to the cache at the write rate — 1.25× the input price on both providers.
  • OpenAI's GPT-6 Astra and GPT-6.1 Sol bill any request over 272K input tokens at 2× input and cache rates and 1.5× output. Anthropic charges the same rate across the full 1M window.
  • A session stops when its input exceeds the model's context window. Real agents summarise or trim before then; this tool shows you when you'd have to.
  • Excludes batch discounts, fast or priority tiers, data-residency premiums and tool-specific charges such as web search. Tokenizers differ between providers, so equal token counts are not equal text.
  • GPT-6 Luna: Context window and long-context rules not published at time of writing; modelled without either.

Prices from Anthropic's and OpenAI's own pricing documentation, checked 30 September 2026. They change often.

Related