In short: the cost of an AI agent breaks down into three items: development (once), inference (every time it's used) and monitoring (ongoing). Inference can be calculated precisely: number of requests × tokens per request × model price. In the detailed example below, prompt caching cuts the inference bill roughly in half.
The three cost items
- Development. Scoping the use case, a prototype on your real data, integration with your application and tools, test sets to measure quality. It's often the largest item at first, and the only one that doesn't depend on volume.
- Inference. What you pay the model provider (or in hardware, for a local model) every time the agent works.
- Monitoring. Logs, cost tracking, quality checks over time, adjustments when data or usage changes.
Reference prices
Models are billed per million tokens, for input (what you send) and output (what the model generates). Anthropic's public prices as of 10 October 2026, in US dollars, before discounts:
| Model | Input | Output | Cache read |
|---|---|---|---|
| Claude Haiku 4.5 | $1 | $5 | $0.10 |
| Claude Sonnet 5.5 | $2 | $10 | $0.10 |
| Claude Opus 5.5 | $4 | $20 | $0.20 |
Source: Anthropic's pricing page. Other providers (OpenAI, Google, Mistral) publish pricing of the same kind: the reasoning below applies in exactly the same way.
Two discounts matter a lot:
- Prompt caching: the part of the context that doesn't change from one request to the next (instructions, documentation, tool definitions) is read from the cache at a fraction of the normal price. The initial write to the cache costs a little more than normal input (1.25 times for a 5-minute cache).
- Batch processing (Batch API): a 50% discount on input and output, for tasks that can wait a few hours.
A worked example
Take an agent that handles customer requests: it reads the request, checks a few tools (knowledge base, history), then drafts a reply.
Assumptions: 1,000 requests a day, 20,000 input tokens per request (instructions, tools, context and back-and-forth included), 2,000 output tokens, using Claude Sonnet 5.5.
Without optimisation:
- input: 20,000 tokens × $2 / million = $0.04;
- output: 2,000 tokens × $10 / million = $0.02;
- $0.06 per request, i.e. $60 a day, about $1,800 a month (30 days).
With prompt caching, if 15,000 of the 20,000 input tokens are common to every request:
- cached part: 15,000 × $0.10 / million = $0.0015;
- variable part: 5,000 × $2 / million = $0.01;
- output unchanged: $0.02;
- about $0.032 per request, i.e. about $32 a day and about $950 a month, excluding cache writes, which are small when traffic is steady.
Running the simple steps (sorting, extracting, summarising) on a smaller model and keeping the largest one where it makes a difference can cut the bill further. But be careful: a cheaper model that fails one task in five and forces a retry isn't cheaper.
The costs people forget
- Agent loops. An agent that calls several tools resends the whole history at each step: the context grows with every turn. That's often where the bill gets out of hand.
- Retries. Network errors, rate limits, malformed answers: every retry is paid for.
- Quality tests. To know whether a model or prompt change actually helps, you have to rerun test sets, which consume tokens too.
- Human oversight, until the agent has proven itself on real cases.
My method for costing a project
- Measure from the prototype the real number of tokens per task, on your data, instead of estimating it.
- Reason in cost per successful task, not per request.
- Turn on the free optimisations first: prompt caching and batch processing.
- Only then compare models on your own cases, and consider a local model if privacy or volume justifies it.
Development is costed after scoping, once the scope and the data are known. Inference, on the other hand, can be costed to the cent from the prototype onwards: it's the best basis for a decision.