The model cheatsheet

Which AI is best for which job?

There's no single "best" model — only the best one for the task in front of you. Here's what each is actually good at, ranked from real benchmark, human-preference, and price data. Not opinions.

Reasoning Reasoning · Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback) (Anthropic) — 53.4 Reasoning · GPT-6 Astra (max) (OpenAI) — 52.7 MO Reasoning · Kimi K3 (low) (Moonshot) — 48.3 Coding Coding · Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback) (Anthropic) — 81.6 Coding · GPT-6 Astra (max) (OpenAI) — 76.9 Coding · Muse Spark 1.2 (xhigh) (Meta) — 72.2 Agents Agents · Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback) (Anthropic) — 57.9 Agents · GPT-6 Astra (max) (OpenAI) — 51.0 ZA Agents · GLM 5.3 Flash (Z.ai) — 50.9 Chat & writing Chat & writing · claude-fable-5.1-max (Anthropic) — 1,508 Chat & writing · gemini-3.8-flash-high (Google) — 1,495 Chat & writing · muse-spark-1.3-max (Meta) — 1,490 Speed ST Speed · Step 3.7 Flash (StepFun) — 413 tok/s Speed · Gemini 3.5 Flash-Lite (Google) — 403 tok/s Speed · gpt-oss-120b (low) (OpenAI) — 346 tok/s Long context Long context · Llama 4 Scout 17b 128e Instruct Maas (Meta) — 10M Long context · Gemini Exp 1206 (Google) — 2.1M Long context · Grok 4 Fast Reasoning (xAI) — 2M Best value Best value · Gemma 4 E4B (Non-reasoning) (Google) — 217.5 pts/$ Best value · DeepSeek V4 Flash (Reasoning, High Effort) (DeepSeek) — 213.7 pts/$ Best value · Qwen3.5 4B (Non-reasoning) (Qwen) — 180.0 pts/$ Budget Budget · Llama 3.1 8b (Meta) — $0.035/Mtok Budget · Meta Llama 3.2 1B Instruct (Amazon) — $0.05/Mtok Budget · Nemotron 3.5 Lightning 30b A3b (Perplexity) — $0.051/Mtok WHICH AI? MyToken Tracker ranked from real data

The cheatsheet

Best models for each job

Pick the job, get the shortlist. Each list is ranked by the metric that actually matters for that task — and refreshes as new models land.

🧠

Reasoning & hard problems

Deep multi-step thinking, math, analysis

  1. 1 Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback) 53.4
  2. 2 GPT-6 Astra (max) 52.7
  3. 3 MO Kimi K3 (low) 48.3
  4. 4 Gemini 3.5 Flash (medium) 46.7
  5. 5 DeepSeek V4 Pro (Reasoning, Max Effort) 44.3

Ranked by Intelligence Index

💻

Writing & shipping code

Generating, refactoring, and fixing code

  1. 1 Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback) 81.6
  2. 2 GPT-6 Astra (max) 76.9
  3. 3 Muse Spark 1.2 (xhigh) 72.2
  4. 4 MO Kimi K3 (low) 72.0
  5. 5 ZA GLM 5.3 Flash 71.5

Ranked by Coding Index

🤖

Agents & tool use

Autonomous workflows that call tools

  1. 1 Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback) 57.9
  2. 2 GPT-6 Astra (max) 51.0
  3. 3 ZA GLM 5.3 Flash 50.9
  4. 4 Muse Spark 1.2 (xhigh) 43.2
  5. 5 MO Kimi K3 (low) 39.6

Ranked by Agentic Index

💬

General chat & writing

Everyday assistant, drafting, Q&A

  1. 1 claude-fable-5.1-max 1,508
  2. 2 gemini-3.8-flash-high 1,495
  3. 3 muse-spark-1.3-max 1,490
  4. 4 qwen3.8-max 1,481
  5. 5 ZA glm-5.3-max 1,475

Ranked by LMArena (human votes)

Real-time & low latency

Voice, autocomplete, anything live

  1. 1 ST Step 3.7 Flash 413 tok/s
  2. 2 Gemini 3.5 Flash-Lite 403 tok/s
  3. 3 gpt-oss-120b (low) 346 tok/s
  4. 4 Nemotron 3.5 Lightning 293 tok/s
  5. 5 Qwen3.5 Omni Flash 272 tok/s

Ranked by Output tokens/sec

📚

Huge documents & long context

Whole codebases, books, long transcripts

  1. 1 Llama 4 Scout 17b 128e Instruct Maas 10M
  2. 2 Gemini Exp 1206 2.1M
  3. 3 Grok 4 Fast Reasoning 2M
  4. 4 GPT 5.5 1.1M
  5. 5 Zai GLM 5 2 1M

Ranked by Context window

💎

Best bang for the buck

The most intelligence per dollar

  1. 1 Gemma 4 E4B (Non-reasoning) 217.5 pts/$
  2. 2 DeepSeek V4 Flash (Reasoning, High Effort) 213.7 pts/$
  3. 3 Qwen3.5 4B (Non-reasoning) 180.0 pts/$
  4. 4 ZA GLM 5.3 Flash 176.0 pts/$
  5. 5 IB Granite 4.2 3B 173.3 pts/$

Ranked by Intelligence per $/Mtok

🪙

High-volume on a budget

Cheap, good-enough, at scale

  1. 1 Llama 3.1 8b $0.035/Mtok
  2. 2 Meta Llama 3.2 1B Instruct $0.05/Mtok
  3. 3 Nemotron 3.5 Lightning 30b A3b $0.051/Mtok
  4. 4 CO Command R7b 12 2024 $0.066/Mtok
  5. 5 Qwen Turbo $0.088/Mtok

Ranked by Lowest blended $/Mtok

Ranked from live data · updated 1 hour ago. Model & provider names are trademarks of their owners, shown here only to report public benchmark and price data.

No favorites

How the picks are made

📊

Benchmarks, not vibes

Reasoning, coding, agentic, speed and value come from independent Artificial Analysis indices. Chat is the LMArena leaderboard — millions of blind human votes. Context and budget come from the live price catalog.

🔁

Self-updating

Nothing is hand-picked. When a new model tops a benchmark or a price changes, the cheatsheet re-ranks itself on the next daily sync. No stale "best of 2024" lists.

🎯

One axis at a time

A model can win one job and lose another. We rank each category by the single metric that matters for it, so the shortlist is honest about trade-offs.

Quality & speed from Artificial Analysis; human preference from LMArena; prices from the MyTokenTracker catalog. See the full methodology.

Citation

Use this in your work

Open data, free to cite. Pair it with the price-vs-cost breakdown and the full State of AI.

Copy a citation

Free to use and cite under CC BY 4.0. See how this is measured.

APA

Champlin Enterprises. (2026). Which AI for which job — the model cheatsheet (MyTokenTracker) [Data set]. MyTokenTracker. Retrieved September 20, 2026, from https://mytokentracker.io/which-ai

BibTeX
@misc{mytokentracker-which-ai,
  title        = {Which AI for which job — the model cheatsheet (MyTokenTracker)},
  author       = {{Champlin Enterprises}},
  year         = {2026},
  howpublished = {MyTokenTracker, \url{https://mytokentracker.io/which-ai}},
  note         = {Accessed September 20, 2026. Licensed CC BY 4.0.},
  url          = {https://mytokentracker.io/which-ai}
}

Need a fixed point in time? Every day’s data is permanently archived in the open-data repository, so you can cite a specific date by linking that day’s committed file.

Free weekly digest

The best model keeps changing

New models top these lists every few weeks. Get the weekly digest — what moved, what's now best for what, and what it costs. Free, no account.

No spam, no account. One click to leave.