Gemma 4 26B A4B
Gemma 4 26B A4B IT is an instruction-tuned Mixture-of-Experts (MoE) model from Google DeepMind. Despite 25.2B total parameters, only 3.8B activate per token during inference — delivering near-31B quality at a fraction of the compute cost. Supports multimodal input including text, images, and video (up to 60s at 1fps). Features a 256K token context window, native function calling, configurable thinking/reasoning mode, and structured output support. Released under Apache 2.0.
google/gemma-4-26b-a4b-it
- Context
- 262k tokens
- Completion cap
- 16,384
- Tools
- Yes
- JSON
- Yes
- Released
- 2026-04-03
Call it from Geek Hub
Same OpenAI SDK. Change the base URL and the model id.
curl https://api.geekhub.mx/v1/chat/completions \
-H "Authorization: Bearer $GEEKHUB_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "google/gemma-4-26b-a4b-it",
"messages": [{"role": "user", "content": "Hola"}]
}'Get an API keyFrequently asked questions
- What is Gemma 4 26B A4B?
- Gemma 4 26B A4B IT is an instruction-tuned Mixture-of-Experts (MoE) model from Google DeepMind. Despite 25.2B total parameters, only 3.8B activate per token during inference — delivering near-31B quality at a fraction of the compute cost. Supports multimodal input including text, images, and video (up to 60s at 1fps). Features a 256K token context window, native function calling, configurable thinking/reasoning mode, and structured output support. Released under Apache 2.0. Gemma 4 26B A4B runs on the Geek Hub API (OpenAI-compatible). Model id: google/gemma-4-26b-a4b-it.
- Is Gemma 4 26B A4B free?
- No. Input is $0.0763 per 1M tokens and output is $0.3706 per 1M tokens on Geek Hub (markup included).
- What is the context length of Gemma 4 26B A4B?
- Gemma 4 26B A4B has a 262k tokens context window. It supports up to 16,384 completion tokens.
- Does Gemma 4 26B A4B support tool calling and structured outputs?
- Gemma 4 26B A4B accepts tools and tool_choice for function calling. It also supports structured outputs via a JSON schema in response_format.
- When was Gemma 4 26B A4B released?
- Gemma 4 26B A4B was released on 2026-04-03.
More models from Google
Gemini 2.5 FlashFlash with hybrid reasoning. Price went up in the July refresh ($0.15/$0.60 → $0.30/$2.50) reflecting the provider change.
Gemini 2.5 Flash ImageReplaces Imagen 4. Excellent with text embedded in the image (much better than Flux/DALL·E) at the catalog's lowest price.
Gemini 2.5 Flash-LiteCheapest in the catalog with 1M context. Output at $0.40 / 1M. Alternative when you want to minimize cost.
Gemini 2.5 Pro2M context, multimodal, reasoning. Google updated output pricing ($5 → $10) in the July refresh.
Gemini 3.1 Flash-Lite3.1 Flash-Lite. Cheap and fast for classification, extraction, and light chat.
Gemini 3.1 Pro (preview)Preview of the 3.1 Pro flagship. 2M context. Price for ≤200k tokens tier; scales ~2x above that.