Data & Research illustration

Data & Research

Cheapest LLMs by blended price right now

The lowest blended prices are clustered in embeddings and small coder models, with one outlier at $0.001/Mtok blended.

October 1, 2026 · 2 min read · By MyTokenTracker

Back to Blog

The cheapest public text-model pricing right now is dominated by embeddings and small specialist models, not general-purpose chat stacks. The absolute floor is $0.001/Mtok blended, and the next tier jumps to $0.006/Mtok blended, so the cheapest options are extremely compressed at the bottom.

For developers, that means the budget winners are the models that do one job well, especially retrieval and code-focused workloads. If you are optimizing for raw token cost, this list is where the economics get unusually favorable.

fireworks_ai/accounts/fireworks/models/flux-1-dev-controlnet-union$0.001
fireworks-ai-embedding-up-to-150m$0.006
fireworks-ai-embedding-150m-to-350m$0.012
nscale/Qwen/Qwen2.5-Coder-3B-Instruct$0.015
nscale/Qwen/Qwen2.5-Coder-7B-Instruct$0.015
nebius/Qwen/Qwen2.5-Coder-7B$0.015
vercel_ai_gateway/amazon/titan-embed-text-v2$0.015
lambda_ai/llama3.2-11b-vision-instruct$0.018
Cheapest text models by blended $/1M tokens (lower is cheaper).

The pricing spread is telling. After the single cheapest entry at $0.001/Mtok blended, the next lowest prices are $0.006/Mtok blended and $0.012/Mtok blended, then a cluster at $0.015/Mtok blended. That shape suggests the market is rewarding narrow, efficient models more than broad generalists.

Embeddings are especially cheap here, with two Fireworks entries at $0.006/Mtok blended and $0.012/Mtok blended, plus another embedding option at $0.015/Mtok blended. For retrieval-heavy apps, semantic search, RAG pipelines, and ranking layers, that is a strong signal that the infrastructure cost of tokenizing and embedding content is no longer the expensive part.

The coder models are also priced aggressively, with Qwen variants at $0.015/Mtok blended. That makes them attractive for autocomplete, code transformation, and agent steps where you do not need the broadest reasoning model. The lone vision model in the set sits at $0.018/Mtok blended, which is still low enough to keep multimodal features from dominating budget.

  • The cheapest public text-model pricing is concentrated in embeddings and small specialist models.
  • $0.015/Mtok blended is the main price cluster for coder models, so budget planning can treat that as a practical floor for many code tasks.
  • Retrieval and embedding workloads are now extremely cheap, which shifts spend toward orchestration and product logic instead of tokens.
  • The vision entry at $0.018/Mtok blended stays in the same low-cost band, so multimodal features are not automatically expensive.

About this data. Figures are from MyTokenTracker live data (model prices, value for money, open data API), free to use and cite under CC BY 4.0. Snapshot: October 1, 2026.