Qwen3 TTS (10 languages, 9 speakers)
Default speech synthesis - 9 voices, 10 languages, styled by instruction
Overview
Qwen3-TTS-12Hz-1.7B-CustomVoice is the instruction-controlled member of the Qwen3-TTS family. It ships 9 premium timbres covering combinations of gender, age, language and dialect, and lets you steer timbre, emotion and prosody with natural-language instructions. Its 12Hz tokenizer and discrete multi-codebook architecture support streaming generation across Chinese, English, Japanese, Korean, German, French, Russian, Portuguese, Spanish and Italian.
Strengths
- Tone and emotion driven by plain-language instructions
- Streaming and non-streaming from one model
- 10 languages with dialect voice profiles
Use cases
- Voice agents needing a consistent set of voices
- Narration with per-passage emotional direction
Good to know
This alias serves the CustomVoice checkpoint: 9 fixed timbres steered by instructions. Voice cloning and free-form voice design are separate checkpoints; the function SDK's qwen3_tts helper exposes both inside workflows.
API access
One OpenAI-compatible API for every model in the catalog, plus the CLI and the function SDK inside workflows. Sign up free, add a card for $5 in credits, and these requests work:
# Text-to-speech models are called through the # platform APIs and workflows; list what is deployed: curl https://gateway.graphn.ai/v1/models \ -H "Authorization: Bearer $GRAPHN_API_KEY"
graphn model get qwen3-tts
from foundry_helpers import qwen3_tts
@function
async def speak(text: str) -> dict:
return await qwen3_tts.generate_speech(
text=text,
language="English",
speaker="Aiden",
)