Reasoning models: thinking and chain-of-thought
Reasoning models spend more thinking tokens in exchange for fewer errors on hard tasks. Ideal for math, multi-step planning, deep code review, and agents that should not improvise.
Models in this collection
| Model | Type | Context | Price |
|---|---|---|---|
| Chat | 200k tokens | $1.1991/M in · $4.7963/M out | |
| Chat | 1M tokens | $1.3626/M in · $2.7251/M out | |
| Chat | 1M tokens | $10.9006/M in · $54.503/M out | |
| Chat | 1M tokens | $5.4503/M in · $27.2515/M out | |
| Chat | 1M tokens | $0.4742/M in · $0.9484/M out | |
| Chat | 2M tokens | $1.3626/M in · $10.9006/M out | |
| Chat | 400k tokens | $5.4503/M in · $32.7018/M out | |
| Chat | 1M tokens | $3.2702/M in · $16.3509/M out |
Why these models
o3-mini and Grok Reasoning are tuned for thinking; Opus/Fable and DeepSeek Pro for agentic reasoning; Gemini Pro and GPT-5.6-sol as high-ceiling alternatives.
Use them with Geek Hub
One OpenAI-compatible base URL and API key. Swap any model id from this list without rewriting your client.
Get an API keyFAQ
- Why are they more expensive?
- They often emit internal reasoning tokens. Measure cost per solved task, not only final answer tokens.
- When should I NOT use reasoning?
- Simple classification, short chat, trivial extraction: Flash/mini/Haiku. Reasoning shines when single-shot fails.
- Good for coding?
- Yes for hard bugs and design. For autocomplete and quick refactors, see Models for Coding.