Overview
Qwen3-ASR-1.7B builds on the audio understanding of the Qwen3-Omni foundation model to do language identification and transcription for 30 languages and 22 Chinese dialects, plus English accents from several countries and regions. A single checkpoint serves both streaming and offline inference and transcribes long audio, and Qwen reports it staying robust in complex acoustic environments including singing voice and songs over background music.
Strengths
- Streaming and offline from one checkpoint
- 30 languages plus 22 Chinese dialects
- Robust on noisy audio and singing
Use cases
- Real-time transcription for voice agents
- Batch transcription of recordings
- Multilingual or dialect-heavy audio
Good to know
The "52 languages" in the display name counts 30 languages plus 22 Chinese dialects.
API access
One OpenAI-compatible API for every model in the catalog, plus the CLI and the function SDK inside workflows. Sign up free, add a card for $5 in credits, and these requests work:
# Speech-to-text models are called through the # platform APIs and workflows; list what is deployed: curl https://gateway.graphn.ai/v1/models \ -H "Authorization: Bearer $GRAPHN_API_KEY"
graphn model get qwen3-asr
from foundry_helpers import qwen3_asr
@function
async def transcribe(path: str) -> str:
result = await qwen3_asr.transcribe(path, language="en", return_timestamps=True)
return result["text"]