How LLM API pricing works
OpenAI, Anthropic, Google and the other model makers charge per token, with separate prices for the tokens you send (input) and the tokens the model writes (output), usually quoted per million tokens. A request costs:
Multiply by the number of requests and you have the bill. A chatbot that sends a 1,000-token prompt with history and gets a 500-token answer, 1,000 times a day, costs at three typical price levels:
| Price level | Input / output per 1M tokens | Per month |
|---|---|---|
| Budget model | $0.1 / $0.4 | $9 |
| Mid-range model | $1 / $4 | $90 |
| Top model | $5 / $25 | $525 |
The gap between the cheapest and the most capable models is often a factor of 50 or more. Many teams route simple tasks to a small model and keep the big GPT, Claude or Gemini model for the hard ones. Open models like Llama, Qwen or DeepSeek are among the cheapest through hosted APIs.
Ways to lower the cost
- Shorter answers: output tokens cost the most, so ask for concise replies and set a maximum length.
- Prompt caching: long system prompts and documents that repeat are billed at a fraction of the input price when served from the cache.
- Batch processing: many providers offer about half price for requests that can wait a few hours.
- Trim the context: chat history grows with every turn; summarize or cut old messages.
How to use the AI token cost calculator
- 1Enter how many tokens a typical request sends and receives, or paste a sample text to estimate it, and how many requests you expect per day or month.
- 2Compare the models in the table, filter by provider or search for one, and add a custom price if yours is different.
- 3Copy the comparison for your budget, or share the link with your team.
Frequently asked questions
What is a token?
Language models read and write text in small pieces called tokens. In English one token is about four characters or three quarters of a word, so 1,000 tokens are roughly 750 words. Other languages, code and numbers often need more tokens for the same length.
Why is output more expensive than input?
Generating text takes the model far more work per token than reading it, so providers charge output tokens at two to eight times the input price. Long answers therefore drive the bill more than long prompts.
Where do the prices come from?
From the public model list of OpenRouter, refreshed once a day, in US dollars per million tokens. For the big providers these match the list prices of OpenAI, Anthropic, Google and others in most cases, but discounts, regional prices, batch rates and long-context surcharges can differ. Check the provider’s own pricing page before you commit to a budget.
What does cached input mean?
Many APIs charge less for input that repeats across requests, like a long system prompt or a document you ask several questions about. If you use prompt caching, enter the share of input that is served from the cache and the calculator applies the cache price where the provider publishes one.
Can I add a model or my own price?
Yes. Use the custom model row to enter any input and output price per million tokens, for example a negotiated rate or a model that is not in the list. It is compared with the others in the same table.