Reasoning models: thinking and chain-of-thought

Reasoning models spend more thinking tokens in exchange for fewer errors on hard tasks. Ideal for math, multi-step planning, deep code review, and agents that should not improvise.

8 modelsUpdated 2026-08-27

Models in this collection

ModelTypeContextPrice
o3-mini (reasoning)OpenAIChat200k tokens$1.1991/M in · $4.7963/M out
Grok 4.20 (reasoning)xAIChat1M tokens$1.3626/M in · $2.7251/M out
Claude Fable 5AnthropicChat1M tokens$10.9006/M in · $54.503/M out
Claude Opus 5AnthropicChat1M tokens$5.4503/M in · $27.2515/M out
DeepSeek V4 ProDeepSeekChat1M tokens$0.4742/M in · $0.9484/M out
Gemini 2.5 ProGoogleChat2M tokens$1.3626/M in · $10.9006/M out
GPT-5.6 SolOpenAIChat400k tokens$5.4503/M in · $32.7018/M out
Claude Sonnet 5AnthropicChat1M tokens$3.2702/M in · $16.3509/M out

Why these models

o3-mini and Grok Reasoning are tuned for thinking; Opus/Fable and DeepSeek Pro for agentic reasoning; Gemini Pro and GPT-5.6-sol as high-ceiling alternatives.

Use them with Geek Hub

One OpenAI-compatible base URL and API key. Swap any model id from this list without rewriting your client.

Get an API key

FAQ

Why are they more expensive?
They often emit internal reasoning tokens. Measure cost per solved task, not only final answer tokens.
When should I NOT use reasoning?
Simple classification, short chat, trivial extraction: Flash/mini/Haiku. Reasoning shines when single-shot fails.
Good for coding?
Yes for hard bugs and design. For autocomplete and quick refactors, see Models for Coding.

More collections