Overview
Qwen3-Next-80B-A3B-Instruct is a high-sparsity mixture-of-experts model: 80B total parameters with only 3B active per token, using hybrid Gated DeltaNet plus Gated Attention in place of standard attention. Qwen reports it performing on par with the 235B model on some benchmarks while handling ultra-long context notably better, which is why it is the default chat model here. It is an instruct-only checkpoint and never emits reasoning blocks.
Strengths
- Only 3B parameters active per token
- Holds accuracy at very long context
- Reliable tool calling
Use cases
- Default assistant and agent backbone
- High-volume or cost-sensitive workloads
- Long-document question answering
Good to know
Instruct-only, so there is no thinking mode. Served at 131K context although the weights natively support 262K.
API access
One OpenAI-compatible API for every model in the catalog, plus the CLI and the function SDK inside workflows. Sign up free, add a card for $5 in credits, and these requests work:
curl https://gateway.graphn.ai/v1/chat/completions \
-H "Authorization: Bearer $GRAPHN_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "qwen3-80b",
"messages": [{"role": "user", "content": "Hello"}]
}'graphn model chat qwen3-80b -m "What is GraphN?"
@function
async def ask(question: str) -> str:
return await chat.complete(question, model="qwen3-80b")