An AI agent doesn't make one model call per task. It plans, calls a tool, reads the result, and calls the model again, often many times. Each call re-sends everything gathered so far, so the cost of a task grows much faster than its number of steps. This calculator models that loop, so you can price an agent feature before real traffic arrives.
Figures checked by Muhammad Ahmad against Claude Platform Docs: Pricing, OpenAI API pricing and Gemini API pricing. Prices and rules change, so confirm with the official source before relying on them. How we check the math
How it works
Every step sends the fixed prompt (system instructions and tool definitions) plus the context built up by earlier steps: each earlier step's output and the tool results it produced.
For n steps, total input = n × prompt + (output + tool results) × n(n − 1) ÷ 2. The second term is the accumulated context, and it grows with the square of the number of steps. Total output = n × output per step.
Cost per run = input tokens ÷ 1,000,000 × the model's input price + output tokens ÷ 1,000,000 × its output price. Retries multiply it: a 10% retry rate makes the average task cost 1.1 runs.
Prices come from the same verified dataset as the AI API Cost Calculator. The calculator uses standard rates, so prompt caching, which many providers apply to the repeated prompt, would lower the real bill.
A worked example
An 8-step agent with a 3,000-token prompt, 400 output tokens and 800 tokens of tool results per step sends 57,600 input tokens and 3,200 output tokens per run. The final step alone carries 11,400 input tokens. On Claude Sonnet 5 ($2 in, $10 out per million) that is $0.1472 per run, or $0.1619 per task with 10% retries: $161.92 a month for 1,000 tasks. More than half of the input, 58.3%, is accumulated context rather than the fixed prompt.
Questions people ask
Why do AI agents cost more than chatbots?
Agents make many model calls per task and feed tool results back in, so each call carries more context than the last. A task with 8 steps can easily use ten times the tokens of a single chat reply.
How can I reduce AI agent costs?
Cache the fixed prompt and tool definitions, trim or summarize tool results before feeding them back, cap the number of steps, and route simple steps to a cheaper model.
Are thinking or reasoning tokens included?
Put them in output per step. Providers bill reasoning tokens as output, and on long tasks they can be the largest part of the output bill.
Does prompt caching change the result?
Yes. Cached input is billed at a fraction of the standard rate by most providers, and agents repeat the same prompt on every step. Treat this result as an upper estimate if you use caching.
Can a more expensive model be cheaper for agents?
Sometimes. A stronger model that finishes in fewer steps, or fails less often, can cost less per task than a cheap model that loops. Compare cost per completed task, not price per token.