Back
All Qwen models
TEXT → AUDIO • CURATED • UPDATED JUL 21, 2026

Qwen-Audio-3.0-TTS

Hosted Alibaba TTS across 16 languages, in a fast tier and a fidelity tier

Qwen-Audio-3.0-TTS is Alibaba Tongyi Lab's hosted text-to-speech model, released 21 July 2026 and served through Alibaba Cloud Model Studio rather than as downloadable weights. It ships in two tiers: Flash, tuned for real-time interaction at roughly 300ms first-packet latency, and Plus, tuned for high-quality generation where naturalness and timbre fidelity matter more than speed. It covers 16 languages — Arabic, Chinese, English, French, German, Indonesian, Italian, Japanese, Korean, Malay, Portuguese, Russian, Spanish, Tagalog, Thai and Vietnamese — and improves fidelity on Chinese dialects over the previous generation.

Pricing Paid
Platforms api
API Yes
Open Source No
Modalities Text → Audio
Best For Best for Multilingual Speech
Date Added 2026-07-21

The two-tier split is the useful part: most TTS vendors make you pick latency or fidelity across the whole account, and this exposes both behind one API. The 300ms-level first-packet latency on the Flash tier is fast enough for live conversational use.

Call it through Alibaba Cloud Model Studio — there are no weights to download. Pick the Flash tier for live conversational work where first-packet latency governs the experience, and the Plus tier for narration or dubbing where listeners will notice timbre. Check current per-character pricing in Model Studio before committing to volume.

1 Choose the tier per workload, not per account — the API exposes both
2 Use Flash for anything a user waits on, Plus for anything they replay
3 Test your target dialect specifically, since dialect fidelity was the headline change
4 Stream output for long passages instead of waiting on full synthesis
5 Confirm pricing in Model Studio, as the tiers are billed differently
Announcement Model Studio
NotebookLM Suno ElevenLabs Resemble AI Gemini Omni

Live Multilingual Support Agent

Give a voice agent low-latency speech in a caller's own language.

STEPS:
  1. Detect the caller's language from the opening turn
  2. Route synthesis to the Flash tier for responsiveness
  3. Stream audio out as the reply is generated
  4. Fall back to text if latency budget is exceeded
  5. Log per-language quality for review

Localising a Video Library

Re-voice an existing catalogue across the supported languages.

STEPS:
  1. Export the source script per segment
  2. Synthesise with the Plus tier for fidelity
  3. Spot-check timbre and pacing per language
  4. Align audio to the original timings
  5. Publish with per-language audio tracks
Paid

Requires a paid subscription.

📚

What is Text-to-Audio AI? Complete Guide 2026

Text-to-audio AI generates voice, music, and sound effects from text descriptions. How AI audio tool...

What is AI Voice Generation? Complete Guide 2026

AI voice generation creates natural-sounding speech from text. How AI voice synthesis tools generate...

What is AI Music Generation? Complete Guide 2026

AI music generation creates original music from text descriptions. How AI music tools compose melodi...

How to Use Text-to-Audio AI Tools: Complete Guide 2026

Text-to-audio AI tools for music, voice, and sound generation. Prompt engineering, model capabilitie...

AI Music Production Tools: What Producers Actually Use

Discover the best AI tools for music production: Suno, Stable Audio 2.5, and other AI music generato...

View Qwen-Audio-3.0-TTS Alternatives (2026) →

Compare Qwen-Audio-3.0-TTS with 5+ similar text → audio AI tools.

Q

Is Qwen-Audio-3.0-TTS free?

A

No, Qwen-Audio-3.0-TTS requires a paid subscription.

Q

Does Qwen-Audio-3.0-TTS have an API?

A

Yes, Qwen-Audio-3.0-TTS offers an API for programmatic integration.

Q

What is Qwen-Audio-3.0-TTS best for?

A

Qwen-Audio-3.0-TTS is best for Multilingual Speech. Qwen-Audio-3. The two-tier split is the useful part: most TTS vendors make you pick latency or fidelity across the whole account, and this exposes both behind one API. The 300ms-level first-packet latency on the Flash tier is fast enough for live conversational use.

Q

What platforms does Qwen-Audio-3.0-TTS support?

A

Qwen-Audio-3.0-TTS supports api.

Q

Is Qwen-Audio-3.0-TTS open source?

A

No, Qwen-Audio-3.0-TTS is not open source.

Q

How do I get started with Qwen-Audio-3.0-TTS?

A

Call it through Alibaba Cloud Model Studio — there are no weights to download. Pick the Flash tier for live conversational work where first-packet latency governs the experience, and the Plus tier for narration or dubbing where listeners will notice timbre. Check current per-character pricing in Mod...

Q

How do I generate audio with Qwen-Audio-3.0-TTS?

A

Qwen-Audio-3.0-TTS creates audio from text descriptions. Enter prompts describing the type of audio you want (voice, music, sound effects) along with style, tone, and duration details.

🏷️

Work on Qwen-Audio-3.0-TTS? You're hand-reviewed in our directory. Add this badge to your site — it links back to this profile.

Featured on CuratedAI