← Managed inference

Qwen3 TTS (10 languages, 9 speakers)

Default speech synthesis - 9 voices, 10 languages, styled by instruction

Overview

Qwen3-TTS-12Hz-1.7B-CustomVoice is the instruction-controlled member of the Qwen3-TTS family. It ships 9 premium timbres covering combinations of gender, age, language and dialect, and lets you steer timbre, emotion and prosody with natural-language instructions. Its 12Hz tokenizer and discrete multi-codebook architecture support streaming generation across Chinese, English, Japanese, Korean, German, French, Russian, Portuguese, Spanish and Italian.

Strengths

  • Tone and emotion driven by plain-language instructions
  • Streaming and non-streaming from one model
  • 10 languages with dialect voice profiles

Use cases

  • Voice agents needing a consistent set of voices
  • Narration with per-passage emotional direction

Good to know

This alias serves the CustomVoice checkpoint: 9 fixed timbres steered by instructions. Voice cloning and free-form voice design are separate checkpoints; the function SDK's qwen3_tts helper exposes both inside workflows.

API access

One OpenAI-compatible API for every model in the catalog, plus the CLI and the function SDK inside workflows. Sign up free, add a card for $5 in credits, and these requests work:

qwen3-tts · curl
# Text-to-speech models are called through the
# platform APIs and workflows; list what is deployed:
curl https://gateway.graphn.ai/v1/models \
  -H "Authorization: Bearer $GRAPHN_API_KEY"