Token Insights illustration

Token Insights

Reasoning Tokens vs Output Tokens: What Costs More

Reasoning tokens are often billed like output tokens, so they can quietly dominate your bill. Here’s how to spot the expensive part.

October 4, 2026 · 6 min read · By MyTokenTracker

Back to Blog

Reasoning tokens usually cost more than input tokens, and in many stacks they are billed at the same rate as output tokens. That means the expensive part of an AI workflow is often not the prompt, it is the model thinking and talking back.

Why reasoning tokens matter more than most devs expect

If you only look at prompt size, you miss the part of the bill that tends to grow fastest. A coding agent can send a modest input, then spend a lot of tokens on internal reasoning, tool planning, and final output. The result is a bill that looks small at the start and larger at the end.

For developers, the practical rule is simple: input tokens are the easy part to estimate, output and reasoning tokens are the part that can surprise you. That is especially true in agentic workflows where the model loops, retries, and explains itself in long responses.

When you want to compare models or platforms, don’t just ask what the prompt costs. Ask what the full task costs, including the tokens the model burns while it thinks. If you want a broader market view, the AI Cost Index is the cleanest place to see blended pricing at a glance.

Token math: a worked example that shows the trap

Let’s use one concrete example. Suppose a coding task consumes 20,000 input tokens, 8,000 output tokens, and 12,000 reasoning tokens. If reasoning tokens are billed like output tokens, then the bill is driven by 20,000 input tokens plus 20,000 output-class tokens.

Input cost = 20,000 / 1,000,000 × input price
Output-class cost = 20,000 / 1,000,000 × output price

Now plug in a few real model prices.

  • gpt-4o: input is $2.5 per 1M tokens, output is $10 per 1M tokens.
  • claude-sonnet-4-5: input is $3 per 1M tokens, output is $15 per 1M tokens.
  • deepseek-chat: input is $0.28 per 1M tokens, output is $0.42 per 1M tokens.

For gpt-4o, the arithmetic is:

20,000 × $2.5 / 1,000,000 = $0.05
20,000 × $10 / 1,000,000 = $0.20
Total = $0.25

For claude-sonnet-4-5:

20,000 × $3 / 1,000,000 = $0.06
20,000 × $15 / 1,000,000 = $0.30
Total = $0.36

For deepseek-chat:

20,000 × $0.28 / 1,000,000 = $0.0056
20,000 × $0.42 / 1,000,000 = $0.0084
Total = $0.014

The lesson is not that one model is always better. The lesson is that once reasoning tokens pile up, output-class pricing dominates the bill very quickly. A model with cheap input pricing can still get expensive if it emits a lot of reasoning and final text.

How much does the output side really dominate?

Here’s a small comparison using only the real prices we know. The key is the gap between input and output pricing, because reasoning tokens usually land on the output side of that gap.

Model Input per 1M Output per 1M Output vs input
gpt-4o-mini $0.15 $0.6 4x
deepseek-chat $0.28 $0.42 1.5x
gemini-2.5-flash $0.3 $2.5 8.3x
gpt-4o $2.5 $10 4x
claude-sonnet-4-5 $3 $15 5x
claude-opus-4-1 $15 $75 5x

That table tells you why reasoning-heavy workflows can feel cheap at the prompt layer and expensive at the task layer. If the model spends most of its token budget on output-class tokens, your bill tracks the output rate much more than the input rate.

For a market-wide view of how these prices stack up, check the live model price list. If you want a blended benchmark instead of raw per-model pricing, the frontier and budget baskets in the AI Cost Index are useful shorthand.

Where developers get blindsided

Reasoning costs hide in a few common places:

  • Long tool loops, where the agent thinks, calls tools, then explains the result.
  • Verbose answers, where the model writes more than you asked for.
  • Retry chains, where failed attempts still burn output-class tokens.
  • Large refactors, where the model needs more context and more explanation.

That is why a task can look efficient in a demo and expensive in production. The prompt is only the first hop. The bill is shaped by everything after that, especially if the model is allowed to reason out loud or produce long intermediate steps.

If you care about whether a model is actually worth the spend, the right question is not just “what does it cost?” It is “what do I get per dollar?” That is exactly the kind of question the value-for-money view is meant to answer.

How to reduce reasoning-driven spend without guessing

You do not need to stop using reasoning models. You need to make the reasoning work for you instead of against you.

  1. Keep prompts tight, because extra context can trigger extra explanation.
  2. Cap response length when you only need a patch, summary, or decision.
  3. Use cheaper models for boring steps, then escalate only when needed.
  4. Track success rate, because a cheap failed task is still wasted spend.
  5. Watch the output share, since that is where reasoning usually lands.

There is also a pricing angle. The AI Cost Index shows a blended frontier basket at $4.64 per 1M tokens and a budget basket at $0.74 per 1M tokens, both using a 3:1 input:output blend. That does not tell you everything, but it gives you a fast sense of how much the market is charging for typical usage patterns.

For teams that want the raw numbers behind the index, the open datasets are published under CC BY 4.0 in the open data section.

How to track this in your own stack

The easiest way to see reasoning costs is to measure what your coding agent already logs locally. Run npx mytokentracker and it prints a 30-day cost table at API prices without sending usage data off your machine. If you want a dashboard, create a free account and run npx mytokentracker init, which uploads history and then syncs every 30 minutes.

MyTokenTracker captures cost, input, output, cache, and reasoning tokens, plus latency and success, broken down by provider, model, platform, and use-case. That is the data you need if you want to answer questions like, “Did this agent spend more on thinking than on doing?”

It works with Claude Code, Codex, GitHub Copilot CLI, Antigravity, OpenCode, Amp, Qwen Code, and more. For supported wrappers, it also auto-captures OpenAI, Anthropic, Gemini, and Mistral, and for anything else you can send a single POST to the events API.

Most importantly, only daily totals per agent and model are uploaded. Prompts, code, and file names stay local.

FAQ

Are reasoning tokens always billed separately?

Not always as a separate line item, but in practice they often behave like output tokens for billing. The exact accounting depends on the provider and model, so the safe assumption is that reasoning is part of the expensive side of the meter unless the provider says otherwise.

Which is worse for cost, long prompts or heavy reasoning?

It depends on the task, but heavy reasoning usually hurts more because it can multiply output-class usage after the prompt is already sent. A long prompt is a one-time hit. Repeated reasoning, retries, and verbose answers can keep adding cost.

What’s the fastest way to see if reasoning is eating my budget?

Track input, output, reasoning, and success together for a few days. If output-class tokens are a large share of total spend, or if failed tasks are still expensive, reasoning is probably the main driver. A local tracker like MyTokenTracker makes that visible without changing how you work.

If you want to measure the real cost of your own agents, install MyTokenTracker and check the numbers on your next 30 days of work.