Data & Research illustration

Data & Research

Cheapest LLMs by blended price right now

The lowest public blended prices are concentrated in embedding and smaller coder models, with several options at $0.015/Mtok blended.

October 5, 2026 · 2 min read · By MyTokenTracker

Back to Blog

The cheapest public models right now are not general-purpose chat giants, they are embedding models and smaller code-focused models. The floor is $0.001/Mtok blended, and the next tiers stay extremely low through $0.018/Mtok blended.

For developers, that means token cost is no longer the main constraint for lightweight retrieval, indexing, and coding workflows, but model choice still matters for capability.

fireworks_ai/accounts/fireworks/models/flux-1-dev-controlnet-union$0.001
fireworks-ai-embedding-up-to-150m$0.006
fireworks-ai-embedding-150m-to-350m$0.012
nscale/Qwen/Qwen2.5-Coder-3B-Instruct$0.015
nscale/Qwen/Qwen2.5-Coder-7B-Instruct$0.015
nebius/Qwen/Qwen2.5-Coder-7B$0.015
vercel_ai_gateway/amazon/titan-embed-text-v2$0.015
lambda_ai/llama3.2-11b-vision-instruct$0.018
Cheapest text models by blended $/1M tokens (lower is cheaper).

The price spread is tight at the bottom, but the model mix is telling. The absolute cheapest entry, fireworks_ai/accounts/fireworks/models/flux-1-dev-controlnet-union at $0.001/Mtok blended, is followed by fireworks-ai-embedding-up-to-150m at $0.006/Mtok blended and fireworks-ai-embedding-150m-to-350m at $0.012/Mtok blended. That pattern says the lowest-cost segment is dominated by infrastructure-adjacent workloads, not broad reasoning chat.

At $0.015/Mtok blended, several text models cluster together: nscale/Qwen/Qwen2.5-Coder-3B-Instruct, nscale/Qwen/Qwen2.5-Coder-7B-Instruct, nebius/Qwen/Qwen2.5-Coder-7B, and vercel_ai_gateway/amazon/titan-embed-text-v2. That is a useful budget signal, because once you are in this band, the bigger decision is usually quality, latency, and context behavior, not raw token spend.

The last model in the set, lambda_ai/llama3.2-11b-vision-instruct at $0.018/Mtok blended, is still very cheap, but it is already above the coder and embedding cluster. For teams building vibe-coded tools, this means you can keep retrieval and code assistance cheap enough that experimentation is not gated by usage cost, while reserving higher-cost models for cases where they clearly buy better output.

  • Cheapest public price in this set is $0.001/Mtok blended, but it is not a general chat benchmark.
  • Embedding models dominate the lowest-cost tiers, with prices at $0.006/Mtok blended and $0.012/Mtok blended.
  • Several coder models sit at $0.015/Mtok blended, so budget differences there are minimal.
  • At this range, capability and latency matter more than token cost for most developer workflows.

About this data. Figures are from MyTokenTracker live data (model prices, value for money, open data API), free to use and cite under CC BY 4.0. Snapshot: October 5, 2026.