Resources / AI Agent Cost Estimator
Describe one run of your agent and how often it runs. The estimator prices it on nine Claude, GPT and Gemini models at their official list prices, checked on 3 October 2026.
Your agent
Model cost per month
Same workload on every model, per month
The full report adds the cost at 2× and 10× volume, where caching and model choice save the most, and the hosting and monitoring costs to budget next to the model bill.
Prices used
| Provider | Model | Input | Cached input | Output | Note |
|---|---|---|---|---|---|
| Anthropic | Claude Opus 5.5 | $4.00 | $0.200 | $20.00 | |
| Anthropic | Claude Sonnet 5.5 | $2.00 | $0.200 | $10.00 | |
| Anthropic | Claude Haiku 4.5 | $1.00 | $0.100 | $5.00 | |
| Anthropic | Claude Fable 5.1 | $10.00 | $0.250 | $50.00 | |
| OpenAI | gpt-6-astra | $10.00 | $1.000 | $50.00 | |
| OpenAI | gpt-6.1-sol | $2.00 | $0.100 | $10.00 | |
| OpenAI | gpt-6-luna | $0.10 | $0.010 | $0.50 | |
| Gemini 3.1 Pro Preview | $2.00 | $0.200 | $12.00 | Prompts up to 200k tokens | |
| Gemini 3.8 Flash | $0.75 | $0.075 | $3.75 | Rate until 31 Dec 2026; $1.50 / $7.50 from 1 Jan 2027 |
US dollars, standard tier, checked 3 October 2026 on Anthropic, OpenAI and Google pricing pages. Cached input is the cache-read rate; cache writes, batch discounts and regional surcharges are not included.
input tokens × (1 − cached share) × input price + input tokens × cached share × cached priceoutput tokens × output price(input + output per call) × calls per runper run × runs a monthFAQ
From each provider's official API pricing page, checked on 3 October 2026: Anthropic, OpenAI and Google. They are standard list prices in US dollars before any volume discount, batch discount or regional surcharge.
A model call is one request to the model and one answer back. An agent that looks something up, decides, then writes a reply makes a call for each step, and each call sends the conversation so far as input again. That is why input tokens usually outnumber output tokens.
Most providers charge less for input they have seen recently, such as a long system prompt that every run repeats. The cached share is the part of each call's input billed at that lower cached rate.
No. Providers count tokens with different tokenizers, so the same text is a different number of tokens on each model, and real runs vary in length. Use the estimate to compare models and size a budget, then measure real usage in a pilot.
Keep going
We build a pilot on your data, measure the tokens each run actually uses, and give you the monthly cost before you scale it.