Gemma 4 31B
Gemma 4 31B Instruct is Google DeepMind's 30.7B dense multimodal model supporting text and image input with text output. Features a 256K token context window, configurable thinking/reasoning mode, native function calling, and multilingual support across 140+ languages. Strong on coding, reasoning, and document understanding tasks. Apache 2.0 license.
google/gemma-4-31b-it
- Context
- 262k tokens
- Completion cap
- 262,144
- Tools
- Yes
- JSON
- Yes
- Released
- 2026-04-02
Call it from Geek Hub
Same OpenAI SDK. Change the base URL and the model id.
curl https://api.geekhub.mx/v1/chat/completions \
-H "Authorization: Bearer $GEEKHUB_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "google/gemma-4-31b-it",
"messages": [{"role": "user", "content": "Hola"}]
}'Get an API keyFrequently asked questions
- What is Gemma 4 31B?
- Gemma 4 31B Instruct is Google DeepMind's 30.7B dense multimodal model supporting text and image input with text output. Features a 256K token context window, configurable thinking/reasoning mode, native function calling, and multilingual support across 140+ languages. Strong on coding, reasoning, and document understanding tasks. Apache 2.0 license. Gemma 4 31B runs on the Geek Hub API (OpenAI-compatible). Model id: google/gemma-4-31b-it.
- Is Gemma 4 31B free?
- No. Input is $0.109 per 1M tokens and output is $0.3706 per 1M tokens on Geek Hub (markup included).
- What is the context length of Gemma 4 31B?
- Gemma 4 31B has a 262k tokens context window. It supports up to 262,144 completion tokens.
- Does Gemma 4 31B support tool calling and structured outputs?
- Gemma 4 31B accepts tools and tool_choice for function calling. It also supports structured outputs via a JSON schema in response_format.
- When was Gemma 4 31B released?
- Gemma 4 31B was released on 2026-04-02.
More models from Google
Gemini 2.5 FlashFlash with hybrid reasoning. Price went up in the July refresh ($0.15/$0.60 → $0.30/$2.50) reflecting the provider change.
Gemini 2.5 Flash ImageReplaces Imagen 4. Excellent with text embedded in the image (much better than Flux/DALL·E) at the catalog's lowest price.
Gemini 2.5 Flash-LiteCheapest in the catalog with 1M context. Output at $0.40 / 1M. Alternative when you want to minimize cost.
Gemini 2.5 Pro2M context, multimodal, reasoning. Google updated output pricing ($5 → $10) in the July refresh.
Gemini 3.1 Flash-Lite3.1 Flash-Lite. Cheap and fast for classification, extraction, and light chat.
Gemini 3.1 Pro (preview)Preview of the 3.1 Pro flagship. 2M context. Price for ≤200k tokens tier; scales ~2x above that.