Qwen3.5 9B Vision (image+video)
The cheapest way to put images and video in front of a model
Overview
Qwen3.5-9B is a dense 9B checkpoint with a native vision encoder, trained with early fusion over multimodal tokens. It covers 201 languages and dialects and supports 262K native context. At 9B it is the least expensive multimodal option on the platform and is the default for the vision category.
Strengths
- Smallest multimodal model in the catalogue
- 262K native context
- 201 languages and dialects
Use cases
- Bulk image tagging and description
- Screenshot and document triage
- Video question answering where cost dominates
Good to know
Sized for speed and cost. For demanding visual reasoning over dense documents or long video, step up to qwen3.8-27b.
API access
One OpenAI-compatible API for every model in the catalog, plus the CLI and the function SDK inside workflows. Sign up free, add a card for $5 in credits, and these requests work:
curl https://gateway.graphn.ai/v1/chat/completions \
-H "Authorization: Bearer $GRAPHN_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "qwen3-vl",
"messages": [{"role": "user", "content": "Hello"}]
}'graphn model chat qwen3-vl -m "What is GraphN?"
@function
async def describe(image_url: str) -> str:
return await vision.analyze(image_url, "What is in this image?", model="qwen3-vl")