← Managed inference

Qwen3 80B Instruct (131K context)

Everyday chat and agent work at near-flagship quality

Overview

Qwen3-Next-80B-A3B-Instruct is a high-sparsity mixture-of-experts model: 80B total parameters with only 3B active per token, using hybrid Gated DeltaNet plus Gated Attention in place of standard attention. Qwen reports it performing on par with the 235B model on some benchmarks while handling ultra-long context notably better, which is why it is the default chat model here. It is an instruct-only checkpoint and never emits reasoning blocks.

Strengths

  • Only 3B parameters active per token
  • Holds accuracy at very long context
  • Reliable tool calling

Use cases

  • Default assistant and agent backbone
  • High-volume or cost-sensitive workloads
  • Long-document question answering

Good to know

Instruct-only, so there is no thinking mode. Served at 131K context although the weights natively support 262K.

API access

One OpenAI-compatible API for every model in the catalog, plus the CLI and the function SDK inside workflows. Sign up free, add a card for $5 in credits, and these requests work:

qwen3-80b · curl
curl https://gateway.graphn.ai/v1/chat/completions \
  -H "Authorization: Bearer $GRAPHN_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "qwen3-80b",
    "messages": [{"role": "user", "content": "Hello"}]
  }'