Qwen-Audio 3.0 TTS Flash

Model overview
qwen-audio-3.0-tts-flash
Qwen-Audio 3.0 TTS Flash is a speech-synthesis model for low-latency real-time interaction, with time to first audio under 200 ms plus expanded minority-language and Chinese-dialect support, enhanced instruction following, and fine-grained tag control.
audiospeechtext to speechvoice assistantrealtime voiceenterprise access
Representative version
qwen-audio-3.0-tts-flash
| Company | Alibaba / Qwen |
|---|---|
| Release | 2026-07-14 |
| Release type | Model release |
| Parameters | Not disclosed |
| Architecture | Not disclosed |
| Context | Not disclosed |
| Input modalities | text |
| Output modalities | audio |
| Tool use / function calling | Tool use: Not disclosed; Function calling: Not disclosed |
| Structured output / reasoning | Structured output: Not disclosed; Reasoning: Not disclosed |
| Vision / audio / video | Confirmed: Audio |
| Open weights / License | Weight status not disclosed / license not disclosed alibaba-qwen-qwen-audio-3-0-tts-flash-current · License source pending verification |
| Vendor API | Available Alibaba Cloud Model Studio model lifecycle and updatesalibaba-qwen-qwen-audio-3-0-tts-flash-current · 2026-07-14 |
| OpenRouter | Not separately listed on OpenRouter |
| Pricing | Official · checked 2026-07-20 · China mainland (Beijing): text input $0.137521 · audio output no separate charge · per 10K input characters Official · checked 2026-07-20 · Singapore: text input $0.15 · audio output no separate charge · per 10K input characters |
| Benchmark | No reliable public source |
| Public signals | No reliable public source |
| Use cases | Low-latency speech synthesis / Voice cloning / Instruction-controlled narration / voice assistant |
| Claim evidence | Release event / Model specificationsAlibaba Cloud Model Studio model lifecycle and updatesqwen-audio-3.0-tts-flash · 2026-07-14 · Qwen-Audio 3.0 TTS Flash · qwen-audio-3.0-tts-flash |
OpenRouter
Pricing / context / availability
| OpenRouter | Not separately listed on OpenRouter |
|---|---|
| OpenRouter listing | Not separately listed in the current OpenRouter public list |
| Mapped version | qwen-audio-3.0-tts-flash alibaba-qwen-qwen-audio-3-0-tts-flash-current |
| Pricing | Official · checked 2026-07-20 · China mainland (Beijing): text input $0.137521 · audio output no separate charge · per 10K input characters Official · checked 2026-07-20 · Singapore: text input $0.15 · audio output no separate charge · per 10K input characters |
| Context | Not disclosed |
| Speed/Latency | No reliable public source |
| Snapshot time | 2026-08-11T06:28:09.895Z |
| Source | View source |
Verified limitations
Usage boundaries
- Billing is by input characters; audio output is not charged separately, and Beijing and Singapore prices differ.
Missing public information
Public information not yet available
Parameter scale not disclosedNot listed on OpenRouterNo reliable benchmark sourceNo public model-card activity countersNo linkable model-specific reviewSpeed/latency not disclosed