← Managed inference

Qwen3-VL Embeddings (4096 dim, multimodal)

Puts text, images, screenshots and video in one vector space

Overview

Qwen3-VL-Embedding-8B is built on Qwen3-VL-8B-Instruct for multimodal retrieval. It embeds text, images, screenshots, video and arbitrary mixtures of them into a shared space, covers 30+ languages, and emits up to 4096 dimensions with user-selectable output sizes from 64 upward. Context is 32K. Reach for it when a knowledge base has to answer an image query with text documents, or the reverse.

Strengths

  • One shared space for text and images
  • Output dimensions selectable from 64 to 4096
  • Instruction-aware embeddings

Use cases

  • Knowledge bases mixing images and text
  • Screenshot search
  • Cross-modal retrieval such as a text query matching an image

Good to know

At 8B parameters with a 32K context it is much heavier and shorter-context than bge-m3. Choose it only when images or video genuinely need to be in the index.

API access

One OpenAI-compatible API for every model in the catalog, plus the CLI and the function SDK inside workflows. Sign up free, add a card for $5 in credits, and these requests work:

qwen3-vl-embedding · curl
curl https://gateway.graphn.ai/v1/embeddings \
  -H "Authorization: Bearer $GRAPHN_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model": "qwen3-vl-embedding", "input": "Hello"}'