Agent loop cost calculator
An agent resends the entire conversation on every turn, so turn forty costs far more than turn one. Pricing pages quote a rate per token; this shows what that rate adds up to across a session.
Start from a workload
Large system prompt and tool definitions, file contents coming back every turn.
Per session
$2.74
Per month
$2,745
Last turn vs first
3.6×
Turn 40 costs 3.6× as much as turn 1, because every turn resends the whole conversation. Prompt caching cuts this session by 74%.
Cost of each turn — Claude Sonnet 5
Peak $0.108 per turn — hover a bar
Same workload, every model
Monthly cost with caching on. Click a row to chart that model.
| Model | Monthly | Per session |
|---|---|---|
| GPT-6 LunaOpenAI | $137.24 | $0.137 |
| GPT-6.1 SolOpenAI | $2,298 | $2.30 |
| Claude Sonnet 5Anthropic | $2,745 | $2.74 |
| Claude Opus 5.5Anthropic | $4,596 | $4.60 |
| Claude Fable 5.1Anthropic | $10,373 | $10.37 |
| GPT-6 AstraOpenAI | $13,724 | $13.72 |
| Claude Haiku 4.5Anthropic | Context full at turn 35 | — |
The point of this tool: load “Research agent”, pick GPT-6 Astra and watch the bars jump where the conversation crosses 272K tokens. Then switch caching off. The per-token price on the pricing page never changed; the bill did, because an agent is billed for its whole memory on every turn.
Why agent cost grows faster than turn count
Language model APIs are stateless: to continue a conversation you send all of it again. Turn one sends the system prompt and a message. Turn thirty sends the system prompt, twenty-nine earlier exchanges, every tool result that came back, and the new message. The input for each turn grows in a straight line, so the total for a session grows with the square of its length. Doubling the number of turns roughly quadruples the cost. Caching is what makes this affordable — the part of the conversation the model has already seen is billed at a fraction of the normal input rate — which is why the cache-read price matters more for agents than the headline input price does.
How it's calculated
- Input on turn t = system prompt + t × tokens added per turn + (t − 1) × output per turn. Earlier outputs become part of the history.
- With caching on, the cache-hit share of everything already seen is billed at the cache-read rate. Everything else is written to the cache at the write rate — 1.25× the input price on both providers.
- OpenAI's GPT-6 Astra and GPT-6.1 Sol bill any request over 272K input tokens at 2× input and cache rates and 1.5× output. Anthropic charges the same rate across the full 1M window.
- A session stops when its input exceeds the model's context window. Real agents summarise or trim before then; this tool shows you when you'd have to.
- Excludes batch discounts, fast or priority tiers, data-residency premiums and tool-specific charges such as web search. Tokenizers differ between providers, so equal token counts are not equal text.
- GPT-6 Luna: Context window and long-context rules not published at time of writing; modelled without either.
Prices from Anthropic's and OpenAI's own pricing documentation, checked 30 September 2026. They change often.
Related
- Claude vs GPT API pricing — where the bills diverge when the sticker prices don't.
- LLM API cost estimator — for single-shot requests rather than loops.
- Voice-agent cost calculator — the full per-minute bill for phone agents.