← Managed inference

Qwen3.5 9B Vision (image+video)

The cheapest way to put images and video in front of a model

Overview

Qwen3.5-9B is a dense 9B checkpoint with a native vision encoder, trained with early fusion over multimodal tokens. It covers 201 languages and dialects and supports 262K native context. At 9B it is the least expensive multimodal option on the platform and is the default for the vision category.

Strengths

  • Smallest multimodal model in the catalogue
  • 262K native context
  • 201 languages and dialects

Use cases

  • Bulk image tagging and description
  • Screenshot and document triage
  • Video question answering where cost dominates

Good to know

Sized for speed and cost. For demanding visual reasoning over dense documents or long video, step up to qwen3.8-27b.

API access

One OpenAI-compatible API for every model in the catalog, plus the CLI and the function SDK inside workflows. Sign up free, add a card for $5 in credits, and these requests work:

qwen3-vl · curl
curl https://gateway.graphn.ai/v1/chat/completions \
  -H "Authorization: Bearer $GRAPHN_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "qwen3-vl",
    "messages": [{"role": "user", "content": "Hello"}]
  }'