Vision models: multimodal LLMs for images
Vision is not image generation — it is image understanding. Invoices, UI screenshots, diagrams, and product photos go in the prompt. This collection gathers strong multimodals on Geek Hub for visual Q&A and extraction.
Models in this collection
| Model | Type | Context | Price |
|---|---|---|---|
| Chat | 2M tokens | $1.3626/M in · $10.9006/M out | |
| Chat | 1M tokens | $0.327/M in · $2.7251/M out | |
| Chat | 1M tokens | $3.2702/M in · $16.3509/M out | |
| Chat | 272k tokens | $5.4503/M in · $32.7018/M out | |
| Chat | 272k tokens | $0.8175/M in · $4.9053/M out | |
| Chat | 1M tokens | $1.6351/M in · $8.1754/M out | |
| Chat | 1M tokens | $5.4503/M in · $27.2515/M out | |
| Chat | 500k tokens | $2.1801/M in · $6.5404/M out |
Why these models
Gemini leads on visual volume at a good price; Claude and GPT-5.x for finer analysis; Flash/mini when throughput matters more than the last detail.
Use them with Geek Hub
One OpenAI-compatible base URL and API key. Swap any model id from this list without rewriting your client.
Get an API keyFAQ
- Is vision the same as image generation?
- No. Vision = image input. Generation = image output (see Image Models).
- Can I send PDFs?
- Depends on the provider and how you serialize input (page images vs text). Check caps on the model page.
- What’s cheap for vision at scale?
- Gemini Flash or GPT-5.4-mini are usually the cost/quality sweet spot for high volume.