Fast AI models: low latency

Sometimes the metric is milliseconds, not the leaderboard. Autocomplete, hot-path classification, and UIs that cannot wait. This collection prioritizes flash/lite/mini/highspeed variants.

9 modelsUpdated 2026-08-27

Models in this collection

ModelTypeContextPrice
Gemini 2.5 Flash-LiteGoogleChat1M tokens$0.109/M in · $0.436/M out
Gemini 3.5 Flash-LiteGoogleChat1M tokens$0.327/M in · $2.7251/M out
Gemini 3.1 Flash-LiteGoogleChat1M tokens$0.2725/M in · $1.6351/M out
DeepSeek V4 FlashDeepSeekChat1M tokens$0.1526/M in · $0.3052/M out
GPT-4.1 mini (legacy)OpenAIChat1M tokens$0.436/M in · $1.7441/M out
Claude Haiku 4.5AnthropicChat200k tokens$1.0901/M in · $5.4503/M out
Ministral 8BMistralChat128k tokens$0.109/M in · $0.109/M out
Kimi K2.7 Code · HighspeedMoonshot (Kimi)Chat262k tokens$1.0356/M in · $8.7205/M out
Gemini 2.5 FlashGoogleChat1M tokens$0.327/M in · $2.7251/M out

Why these models

Flash Lite and DeepSeek Flash for the latency floor; Haiku and GPT mini when you want speed with more judgment; Kimi Highspeed for concurrent code chat.

Use them with Geek Hub

One OpenAI-compatible base URL and API key. Swap any model id from this list without rewriting your client.

Get an API key

FAQ

Is fast the same as cheap?
They often correlate, not always. Check price and p95 latency in your region with a smoke test.
Good for agents?
As workers, yes. As the only planner, sometimes no. Typical pattern: fast router + strong planner.
Does streaming help?
Yes for perceived speed. Enable stream on the client even when the model is already fast.

More collections