Ling-3.0-flash

*Ling-3.0-flash* is a *124B-parameter Mixture-of-Experts (MoE) model*, with approximately *5.1B parameters activated per token*. The model is designed with *token efficiency and production-scale agentic inference* as key priorities, enabling developers to complete more useful work within constrained token, latency, and serving-cost budgets.

ChatInclusionai262k tokens$0.0229 / $0.0687 · 1M

inclusionai/ling-3.0-flash

Contexto
262k tokens
Máx. completion
32,768
Tools
JSON
Lanzamiento
2026-07-23

Llámalo desde Geek Hub

El mismo SDK de OpenAI. Cambia el base URL y el id del modelo.

curl https://api.geekhub.mx/v1/chat/completions \
  -H "Authorization: Bearer $GEEKHUB_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "inclusionai/ling-3.0-flash",
    "messages": [{"role": "user", "content": "Hola"}]
  }'
Consigue tu API key

Preguntas frecuentes

¿Qué es Ling-3.0-flash?
*Ling-3.0-flash* is a *124B-parameter Mixture-of-Experts (MoE) model*, with approximately *5.1B parameters activated per token*. The model is designed with *token efficiency and production-scale agentic inference* as key priorities, enabling developers to complete more useful work within constrained token, latency, and serving-cost budgets. Ling-3.0-flash corre en el API de Geek Hub (compatible con OpenAI). Id: inclusionai/ling-3.0-flash.
¿Ling-3.0-flash es gratis?
No. El input cuesta $0.0229 / 1M tokens y el output $0.0687 / 1M tokens en Geek Hub (markup incluido).
¿Cuál es el contexto de Ling-3.0-flash?
Ling-3.0-flash tiene una ventana de 262k tokens. Soporta hasta 32,768 tokens de completion.
¿Ling-3.0-flash soporta tool calling y structured outputs?
Ling-3.0-flash acepta tools y tool_choice para function calling. También soporta structured outputs con un JSON schema en response_format.
¿Cuándo se lanzó Ling-3.0-flash?
Ling-3.0-flash se lanzó el 23 de julio de 2026.

Más modelos de Inclusionai