Resources / AI Agent Cost Estimator

AI Agent Cost Estimator: What Each Run Costs, by Model

Describe one run of your agent and how often it runs. The estimator prices it on nine Claude, GPT and Gemini models at their official list prices, checked on 3 October 2026.

Your agent

One run is one task done end to end, such as one ticket answered or one invoice processed.
Each step that asks the model something is one call.
Instructions, context and conversation so far. About 750 words is 1,000 tokens.
Repeated instructions and documents the provider has cached.

Model cost per month

per run
per year
tokens a month
of cost is output

Same workload on every model, per month

The full report adds the cost at 2× and 10× volume, where caching and model choice save the most, and the hosting and monitoring costs to budget next to the model bill.

Thanks. Your report reaches your inbox within one working day, built on the numbers above.

Prices used

List Prices per Million Tokens

ProviderModelInputCached inputOutputNote
AnthropicClaude Opus 5.5$4.00$0.200$20.00
AnthropicClaude Sonnet 5.5$2.00$0.200$10.00
AnthropicClaude Haiku 4.5$1.00$0.100$5.00
AnthropicClaude Fable 5.1$10.00$0.250$50.00
OpenAIgpt-6-astra$10.00$1.000$50.00
OpenAIgpt-6.1-sol$2.00$0.100$10.00
OpenAIgpt-6-luna$0.10$0.010$0.50
GoogleGemini 3.1 Pro Preview$2.00$0.200$12.00Prompts up to 200k tokens
GoogleGemini 3.8 Flash$0.75$0.075$3.75Rate until 31 Dec 2026; $1.50 / $7.50 from 1 Jan 2027

US dollars, standard tier, checked 3 October 2026 on Anthropic, OpenAI and Google pricing pages. Cached input is the cache-read rate; cache writes, batch discounts and regional surcharges are not included.

The formula

  • Input per call: input tokens × (1 − cached share) × input price + input tokens × cached share × cached price
  • Output per call: output tokens × output price
  • Per run: (input + output per call) × calls per run
  • Per month: per run × runs a month

What it leaves out

  • Hosting, monitoring and the workflow platform the agent runs on.
  • Paid tools the agent calls, such as web search.
  • Tokenizer differences: the same text is a different token count on each provider.
  • Retries and failed runs.

FAQ

Questions About This Tool

Where do the model prices come from?

From each provider's official API pricing page, checked on 3 October 2026: Anthropic, OpenAI and Google. They are standard list prices in US dollars before any volume discount, batch discount or regional surcharge.

What is a model call, and why does an agent make several?

A model call is one request to the model and one answer back. An agent that looks something up, decides, then writes a reply makes a call for each step, and each call sends the conversation so far as input again. That is why input tokens usually outnumber output tokens.

What does the cached share do?

Most providers charge less for input they have seen recently, such as a long system prompt that every run repeats. The cached share is the part of each call's input billed at that lower cached rate.

Is the estimate exact?

No. Providers count tokens with different tokenizers, so the same text is a different number of tokens on each model, and real runs vary in length. Use the estimate to compare models and size a budget, then measure real usage in a pilot.

Want the Real Number for Your Agent?

We build a pilot on your data, measure the tokens each run actually uses, and give you the monthly cost before you scale it.

Book a Discovery Call
Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.