Cost Optimization illustration

Cost Optimization

Input vs Output Tokens: What Actually Drives AI Cost

Input tokens usually look cheap until your outputs, retries, and long contexts stack up. Here’s the math developers should actually watch.

October 6, 2026 · 7 min read · By MyTokenTracker

Back to Blog

For most coding workflows, input tokens set the floor, but output tokens often decide the bill. If you want the shortest path to lower AI spend, measure both, then attack the side that your workload actually burns fastest.

Why input and output tokens are not the same cost problem

Developers tend to think in prompts, but pricing is token math. Input tokens are what you send, output tokens are what the model returns, and many tools also track cache and reasoning tokens separately. That matters because the same request shape can swing from cheap to expensive depending on how chatty the model gets, how much context you stuff into the prompt, and whether the task triggers long explanations, code dumps, or retries.

The simplest mental model is this: input tokens are the cost of asking, output tokens are the cost of answering. In agentic workflows, the answer can easily dominate because a single request may produce a plan, a patch, a summary, and a follow-up explanation. If you want a broader market view of where model pricing sits, the AI Cost Index is the quickest way to see blended pricing at a glance.

Here is the key detail that trips people up: LLM prices are quoted per 1M tokens, not per token. So when you compare models, you need to normalize everything to the same unit before you reason about spend.

A worked example with real token math

Say you run a coding assistant task that uses 120,000 input tokens and 30,000 output tokens on gpt-4o. The published price is $2.5 per 1M input tokens and $10 per 1M output tokens.

Input cost  = (120,000 / 1,000,000) * $2.5  = $0.30
Output cost = (30,000 / 1,000,000)  * $10   = $0.30
Total cost  = $0.60

That looks balanced, but the output was only one quarter of the input token count and still cost the same amount. On a model with a bigger output multiplier, the answer side can become the expensive part very quickly.

Now compare that with gpt-4o-mini, priced at $0.15 per 1M input tokens and $0.6 per 1M output tokens.

Input cost  = (120,000 / 1,000,000) * $0.15 = $0.018
Output cost = (30,000 / 1,000,000)  * $0.6  = $0.018
Total cost  = $0.036

Same token counts, very different bill. That is why “cheap model” discussions are incomplete if they ignore the input/output mix.

Which side usually drives spend in coding agents?

In practice, input-heavy workloads and output-heavy workloads fail in different ways. Long system prompts, pasted logs, and huge codebases inflate input cost. Verbose explanations, generated files, refactors, and multi-step reasoning inflate output cost. Agent loops can make both worse because every retry replays the context and asks for another response.

Here is a quick comparison of how a few models price the two sides, using only the current figures we have:

Model Input per 1M Output per 1M What stands out
gpt-4o-mini $0.15 $0.6 Very low on both sides, useful for high-volume utility work
deepseek-chat $0.28 $0.42 Output is only modestly above input, so long responses are less punishing
gemini-2.5-flash $0.3 $2.5 Output is much pricier than input, so verbosity matters
gpt-4o $2.5 $10 Output is 4x input, so generated text can dominate fast
claude-sonnet-4-5 $3 $15 Output is 5x input, so concise prompting pays off

This is also why blended numbers can hide the real shape of your spend. A model can look reasonable on a blended basis, while still being brutal for workflows that generate lots of text. If you want to compare models by actual value, not just raw price, the value-for-money view is the better lens.

How to tell whether input or output is hurting you

Start by looking at the shape of your tasks, not just the total. If you’re sending giant codebases, logs, or multi-file diffs, input is probably the main drag. If you’re asking for full file rewrites, long explanations, or iterative agent responses, output is probably the bigger leak.

  • Input-heavy symptoms: repeated large prompts, lots of pasted context, slow first response, and high spend even when the model replies briefly.
  • Output-heavy symptoms: long answers, generated code blocks, repeated chain-of-thought style verbosity, and costs that spike when tasks get ambiguous.
  • Retry-heavy symptoms: the same job gets re-run with slightly different instructions, which multiplies both sides at once.

For coding agents, the worst pattern is usually not one giant prompt. It is many medium prompts plus many medium outputs, repeated across a session. That is why a tool that can break spend down by provider, model, platform, and use-case is more useful than a raw invoice total.

Want to see what real model availability looks like across the market? Check the live catalog at 4,004 model prices and compare the options your stack is actually using.

How to reduce the side that costs you more

If input is the problem, shrink the context. Strip logs, summarize history, remove duplicate instructions, and avoid sending the same repo state over and over. If output is the problem, tighten the task. Ask for the exact artifact you need, not a lecture around it. Tell the model to be concise when you do not need a wall of text.

There are also model-selection tactics that matter. For example, if your workload is mostly short utility calls, a low-cost model like gpt-4o-mini can keep both sides small. If your workflow needs higher-quality reasoning but the output is still short, a model with a less aggressive output multiplier may be a better fit than one that charges heavily for generated text. The point is not to find the cheapest model in isolation, it is to match the pricing shape to your workload shape.

And if your stack supports caching, use it. Cached input can be a huge deal for repeated system prompts and stable instructions. Even when the prompt content is large, not every token should be paid for repeatedly.

How to track this without guessing

The fastest way to stop guessing is to measure the actual token mix from the tools you already use. MyTokenTracker reads the usage logs coding agents already keep locally, then prints a 30-day cost table at API prices. Nothing leaves your machine unless you choose to sync.

To try it, run npx mytokentracker. That gives you local visibility into cost, input tokens, output tokens, cache, reasoning tokens, latency, and success. If you want a dashboard, create a free account and run npx mytokentracker init, which uploads only daily totals per agent and model, never prompts, code, or file names. It works with Claude Code, Codex, GitHub Copilot CLI, Antigravity, OpenCode, Amp, Qwen Code, and more.

If you are instrumenting your own app, the wrappers auto-capture OpenAI, Anthropic, Gemini, and Mistral. For everything else, a single POST to the events API is enough. The point is to stop treating cost as an after-the-fact surprise and start treating it like any other production metric.

For open data and the broader dataset behind the product, see the public data page.

FAQ

Is output always more expensive than input?

No. It depends on the model. Some models charge much more for output than input, while others keep the gap smaller. What matters is the ratio, not a universal rule.

Why do agentic tools feel more expensive than chat?

Because agents tend to resend context, call the model multiple times, and generate longer outputs. That means you pay for repeated input plus repeated output, not just one neat prompt-response pair.

What should I optimize first, input or output?

Optimize the bigger side in your real workload. If your prompts are huge, trim them. If your outputs are verbose, constrain them. If you do not know which side dominates, track a few days of actual usage before changing anything.

Token spend is rarely one problem. It is usually a mix of too much context, too much output, and too many retries.

If you want to see your own input and output mix instead of guessing, install MyTokenTracker and check the numbers on your next coding session.