Vision models: multimodal LLMs for images

Vision is not image generation — it is image understanding. Invoices, UI screenshots, diagrams, and product photos go in the prompt. This collection gathers strong multimodals on Geek Hub for visual Q&A and extraction.

8 modelsUpdated 2026-08-27

Models in this collection

ModelTypeContextPrice
Gemini 2.5 ProGoogleChat2M tokens$1.3626/M in · $10.9006/M out
Gemini 2.5 FlashGoogleChat1M tokens$0.327/M in · $2.7251/M out
Claude Sonnet 5AnthropicChat1M tokens$3.2702/M in · $16.3509/M out
GPT-5.5OpenAIChat272k tokens$5.4503/M in · $32.7018/M out
GPT-5.4 miniOpenAIChat272k tokens$0.8175/M in · $4.9053/M out
Gemini 3.6 FlashGoogleChat1M tokens$1.6351/M in · $8.1754/M out
Claude Opus 5AnthropicChat1M tokens$5.4503/M in · $27.2515/M out
Grok 4.5xAIChat500k tokens$2.1801/M in · $6.5404/M out

Why these models

Gemini leads on visual volume at a good price; Claude and GPT-5.x for finer analysis; Flash/mini when throughput matters more than the last detail.

Use them with Geek Hub

One OpenAI-compatible base URL and API key. Swap any model id from this list without rewriting your client.

Get an API key

FAQ

Is vision the same as image generation?
No. Vision = image input. Generation = image output (see Image Models).
Can I send PDFs?
Depends on the provider and how you serialize input (page images vs text). Check caps on the model page.
What’s cheap for vision at scale?
Gemini Flash or GPT-5.4-mini are usually the cost/quality sweet spot for high volume.

More collections