Tutorials illustration

Tutorials

What Is Cost per Million Tokens?

Cost per million tokens is the only sane way to compare LLMs. Here’s how to calculate it, read it, and track it in real projects.

October 3, 2026 · 7 min read · By MyTokenTracker

Back to Blog

Cost per million tokens is the cleanest way to compare LLMs because it normalizes pricing across models, providers, and workloads. If you can do one piece of math, do this one, it turns a vague API bill into something you can reason about.

The catch is that most real usage is mixed input and output, so the number that matters is not just the headline price, it is the blended cost for your actual prompt and completion ratio. That’s why the AI Cost Index exists, and why developers should care about the math instead of the marketing.

What cost per million tokens actually means

LLM vendors quote prices per 1,000,000 tokens, not per request and not per character. A token is a chunk of text, and your bill depends on how many tokens you send in and how many the model sends back. The basic formula is simple:

cost = (input_tokens / 1,000,000) * input_price + (output_tokens / 1,000,000) * output_price

That formula is the whole game. If you know your token counts, you can estimate spend before you ship a feature, compare models on equal footing, and spot when a workflow is drifting into expensive territory.

For quick orientation, here are two useful blended reference points from the AI Cost Index:

BasketBlended priceBlended ratio
Frontier$4.64 per 1M tokens3:1 input:output
Budget$0.74 per 1M tokens3:1 input:output

Those baskets are useful because they compress a lot of model variety into one number. If you are comparing toolchains, they give you a rough sense of where your spend will land without pretending every model behaves the same.

How to do the math without fooling yourself

The easiest mistake is to compare only input prices or only output prices. That can be fine for a narrow use case, but most coding agents do both, and some do a lot of both. If your tool spends 80 percent of its tokens on completions, an input-only comparison is basically a lie.

Here is a worked example using only the prices we know. Suppose your app sends 120,000 input tokens and gets back 30,000 output tokens from gpt-4o. The price is $2.5 per 1M input tokens and $10 per 1M output tokens.

Input cost  = 120,000 / 1,000,000 * 2.5  = $0.30
Output cost = 30,000 / 1,000,000 * 10 = $0.30
Total cost = $0.60

Now change nothing except the model and run the same workload on gpt-4o-mini, which is $0.15 per 1M input tokens and $0.6 per 1M output tokens.

Input cost  = 120,000 / 1,000,000 * 0.15 = $0.018
Output cost = 30,000 / 1,000,000 * 0.6 = $0.018
Total cost = $0.036

Same token volume, wildly different bill. That is why cost per million tokens matters more than sticker shock from a monthly invoice. It also explains why cheap-looking prompts can still get expensive when completions are long.

Why blended pricing is better for real developer workflows

Blended pricing answers the question developers actually ask: what will this workflow cost me on average? That matters for code assistants, eval pipelines, agent loops, and anything that chains multiple calls together. A single request is rarely the full story, but a blended rate gives you a practical planning number.

Here is a small comparison using only the specific prices provided in the facts block. These are not universal rankings, just a snapshot of how a few models price their input and output tokens.

ModelInput per 1MOutput per 1M
deepseek-chat$0.28$0.42
gpt-4o-mini$0.15$0.6
gemini-2.5-flash$0.3$2.5
claude-sonnet-4-5$3$15
gpt-4o$2.5$10
claude-opus-4-1$15$75

From a budgeting perspective, the shape of the price matters as much as the absolute number. A model with cheap input but expensive output can be great for retrieval-heavy tasks and painful for long-form generation. A model with balanced pricing can be easier to reason about when you do not know in advance how verbose the response will be.

If you want a broader market view, the live model price catalog is the fastest way to see how many options are out there and how they line up. If you care about whether the market is moving, the State of AI page is the better place to watch the bigger picture.

Where developers get token math wrong

The most common mistake is treating token cost like a static per-request fee. It is not. A request with a tiny prompt and a huge completion can be more expensive than a request with a huge prompt and a tiny completion, even if they look similar at first glance.

Another mistake is ignoring hidden token sources. System prompts, tool schemas, retrieval context, and multi-turn conversation history all count. That means your bill can climb even when the user thinks they asked for “just one small thing.”

There is also a subtle trap in agentic workflows. When a coding agent retries, reflects, or delegates to another model, you can multiply both input and output tokens without noticing. The result is that the apparent cost of one task is actually the sum of several internal calls.

For that reason, the right metric is often not “how much does this model cost?” but “how much does this workflow cost per successful task?” That is the number that tells you whether a feature is economically sane.

How to track this in practice

To see your real cost per million tokens, you need usage logs, not guesses. MyTokenTracker reads the logs coding agents already keep locally, calculates a 30-day cost table at API prices, and keeps the usage data on your machine. If you want a quick check, run npx mytokentracker.

If you want a dashboard, create a free account and run npx mytokentracker init. It asks for the API token, uploads history, then syncs every 30 minutes. It works with Claude Code, Codex, GitHub Copilot CLI, Antigravity, OpenCode, Amp, Qwen Code, and more. Only daily totals per agent and model are uploaded, never prompts, code, or file names.

That matters because the hard part is not the math, it is collecting the right numbers. Once you can see input tokens, output tokens, latency, success, and cost by provider, model, platform, and use-case, you can stop arguing about vibes and start measuring actual spend.

MyTokenTracker also supports OpenAI, Anthropic, Gemini, and Mistral through drop-in wrappers, plus any other provider through a single POST to the events API. If you are instrumenting a product, that is usually enough to get from “we think it is cheap” to “we know exactly what it costs.”

FAQ

Is cost per million tokens enough to compare two models?

It is enough for a first pass, but not the whole story. You also need to know how much output a model tends to generate, whether it supports caching, and how often your workflow retries. Two models with similar blended prices can still produce very different real bills if one is verbose and the other is terse.

Why do some models look cheap on input but expensive overall?

Because output is often priced higher than input, and some tasks produce a lot of output. A model can look attractive for prompt-heavy retrieval work and still become pricey in a generation-heavy workflow. That is why you should always calculate both sides of the formula.

What is the fastest way to estimate my own spend?

Take your typical input tokens, multiply by the input price per 1M, then do the same for output tokens and add them together. If you want the real answer instead of a rough estimate, track your logs directly and compare against actual usage over 30 days.

If you want to see your own numbers instead of averages, install MyTokenTracker and run the local scan.