Estimating the cost of an AI agent by multiplying one model call by the number of steps gives an answer that is far too low. An agent doesn't make independent calls. Each step sends the system prompt and tool definitions again, plus everything earlier steps produced: their outputs and the results of every tool they called. The context grows with every step, and so does the bill.
The math
For n steps with a fixed prompt S, and each step adding its output O and tool results R to the context, total input is n × S + (O + R) × n(n − 1) ÷ 2. The second term, the accumulated context, grows with the square of the number of steps.
Input tokens per run
8 × 3,000 + 1,200 × 8 × 7 ÷ 2equals57,600
Output tokens per run
8 × 400equals3,200
Per task (with retries)
(57,600 ÷ 1M × $2.00 + 3,200 ÷ 1M × $10.00) × 1.1equals$0.1619
Per month
$0.1619 × 1,000 tasksequals$161.92
Each task sends 57,600 input tokens, and 58% of them are accumulated context rather than the fixed prompt. The final step alone carries 11,400 tokens. With 10% of tasks retried, that's about $0.16 a task, or $161.92 a month for 1,000 tasks.
Why more steps cost disproportionately more
Double the agent to 16 steps and the cost per task doesn't double. It rises from about $0.16 to about $0.49, roughly three times as much, because every extra step carries all the context before it. Long-running agents are where budgets break.
Four ways to cut the bill
- Trim tool results before feeding them back. Halving tool output from 800 to 400 tokens a step cuts this example from about $0.16 to about $0.14 a task, and the saving grows with the number of steps.
- Use a cheaper model where it's good enough. The same agent on Claude Haiku 4.5 costs about $0.08 a task, half the Sonnet 5 figure. Many agents route routine steps to a small model and only hard ones to a large one.
- Cache the fixed prompt. System prompts and tool definitions are identical on every step, and most providers bill cached input at a fraction of the normal rate.
- Cap and summarise. Set a maximum number of steps, and replace older context with a short summary once it grows past a threshold.
Measure cost per completed task
The number that matters isn't price per token but cost per task that actually succeeds. A stronger model that finishes in 5 steps can cost less than a cheap one that takes 15, or fails and retries. Log steps, tokens and outcomes per task from the start, and compare models on that.
Tip: Reasoning or thinking tokens are billed as output. On agents that reason at every step, include them in output per step or the estimate will be low.
Questions people ask
- How much does an AI agent cost per task?
- It depends on steps, context and model. An 8-step agent with a 3,000-token prompt on Claude Sonnet 5 costs about $0.16 a task including 10% retries; the same agent at 16 steps costs about $0.49.
- Why do agent costs grow faster than the number of steps?
- Each step re-sends all earlier outputs and tool results, so input grows with the square of the number of steps.
- What is the easiest way to reduce agent costs?
- Trim tool results, cache the fixed prompt and tool definitions, route simple steps to a cheaper model, and cap the number of steps.