Qwen-Audio 3.0 TTS Flash

Alibaba / Qwen
Model overview

qwen-audio-3.0-tts-flash

Qwen-Audio 3.0 TTS Flash is a speech-synthesis model for low-latency real-time interaction, with time to first audio under 200 ms plus expanded minority-language and Chinese-dialect support, enhanced instruction following, and fine-grained tag control.

audiospeechtext to speechvoice assistantrealtime voiceenterprise access
Representative version

qwen-audio-3.0-tts-flash

CompanyAlibaba / Qwen
Release2026-07-14
Release typeModel release
ParametersNot disclosed
ArchitectureNot disclosed
ContextNot disclosed
Input modalitiestext
Output modalitiesaudio
Tool use / function callingTool use: Not disclosed; Function calling: Not disclosed
Structured output / reasoningStructured output: Not disclosed; Reasoning: Not disclosed
Vision / audio / videoConfirmed: Audio
Open weights / LicenseWeight status not disclosed / license not disclosed
alibaba-qwen-qwen-audio-3-0-tts-flash-current · License source pending verification
Vendor APIAvailable
Alibaba Cloud Model Studio model lifecycle and updatesalibaba-qwen-qwen-audio-3-0-tts-flash-current · 2026-07-14
OpenRouterNot separately listed on OpenRouter
PricingOfficial · checked 2026-07-20 · China mainland (Beijing): text input $0.137521 · audio output no separate charge · per 10K input characters
Official · checked 2026-07-20 · Singapore: text input $0.15 · audio output no separate charge · per 10K input characters
BenchmarkNo reliable public source
Public signalsNo reliable public source
Use casesLow-latency speech synthesis / Voice cloning / Instruction-controlled narration / voice assistant
Claim evidenceRelease event / Model specificationsAlibaba Cloud Model Studio model lifecycle and updatesqwen-audio-3.0-tts-flash · 2026-07-14 · Qwen-Audio 3.0 TTS Flash · qwen-audio-3.0-tts-flash
OpenRouter

Pricing / context / availability

OpenRouterNot separately listed on OpenRouter
OpenRouter listingNot separately listed in the current OpenRouter public list
Mapped versionqwen-audio-3.0-tts-flash
alibaba-qwen-qwen-audio-3-0-tts-flash-current
PricingOfficial · checked 2026-07-20 · China mainland (Beijing): text input $0.137521 · audio output no separate charge · per 10K input characters
Official · checked 2026-07-20 · Singapore: text input $0.15 · audio output no separate charge · per 10K input characters
ContextNot disclosed
Speed/LatencyNo reliable public source
Snapshot time2026-08-11T06:28:09.895Z
SourceView source
Verified limitations

Usage boundaries

  • Billing is by input characters; audio output is not charged separately, and Beijing and Singapore prices differ.
Missing public information

Public information not yet available

Parameter scale not disclosedNot listed on OpenRouterNo reliable benchmark sourceNo public model-card activity countersNo linkable model-specific reviewSpeed/latency not disclosed