Qwen3.5-Flash

The Qwen3.5 native vision-language Flash models are built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency. Compared to the 3 series, these models deliver a leap forward in performance for both pure text and multimodal tasks, offering fast response times while balancing inference speed and overall performance.

ChatQwen1M tokens$0.0709 / $0.2834 · 1M

qwen/qwen3.5-flash-02-23

Contexto
1M tokens
Máx. completion
65,536
Tools
JSON
Lanzamiento
2026-02-25

Llámalo desde Geek Hub

El mismo SDK de OpenAI. Cambia el base URL y el id del modelo.

curl https://api.geekhub.mx/v1/chat/completions \
  -H "Authorization: Bearer $GEEKHUB_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "qwen/qwen3.5-flash-02-23",
    "messages": [{"role": "user", "content": "Hola"}]
  }'
Consigue tu API key

Preguntas frecuentes

¿Qué es Qwen3.5-Flash?
The Qwen3.5 native vision-language Flash models are built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency. Compared to the 3 series, these models deliver a leap forward in performance for both pure text and multimodal tasks, offering fast response times while balancing inference speed and overall performance. Qwen3.5-Flash corre en el API de Geek Hub (compatible con OpenAI). Id: qwen/qwen3.5-flash-02-23.
¿Qwen3.5-Flash es gratis?
No. El input cuesta $0.0709 / 1M tokens y el output $0.2834 / 1M tokens en Geek Hub (markup incluido).
¿Cuál es el contexto de Qwen3.5-Flash?
Qwen3.5-Flash tiene una ventana de 1M tokens. Soporta hasta 65,536 tokens de completion.
¿Qwen3.5-Flash soporta tool calling y structured outputs?
Qwen3.5-Flash acepta tools y tool_choice para function calling. También soporta structured outputs con un JSON schema en response_format.
¿Cuándo se lanzó Qwen3.5-Flash?
Qwen3.5-Flash se lanzó el 25 de febrero de 2026.

Más modelos de Qwen