TEXT → AUDIO • CURATED • UPDATED JUL 21, 2026

Qwen-Audio-3.0-TTS

Hosted Alibaba TTS across 16 languages, in a fast tier and a fidelity tier
Price
Paid
API
Yes
Weights
No

Qwen-Audio-3.0-TTS is Alibaba Tongyi Lab's hosted text-to-speech model, released 21 July 2026 and served through Alibaba Cloud Model Studio rather than as downloadable weights. It ships in two tiers: Flash, tuned for real-time interaction at roughly 300ms first-packet latency, and Plus, tuned for high-quality generation where naturalness and timbre fidelity matter more than speed. It covers 16 languages — Arabic, Chinese, English, French, German, Indonesian, Italian, Japanese, Korean, Malay, Portuguese, Russian, Spanish, Tagalog, Thai and Vietnamese — and improves fidelity on Chinese dialects over the previous generation.

The two-tier split is the useful part: most TTS vendors make you pick latency or fidelity across the whole account, and this exposes both behind one API. The 300ms-level first-packet latency on the Flash tier is fast enough for live conversational use.

Pricing Paid
Platforms api
API Yes
Open Source No
Modalities Text → Audio
Best For Best for Multilingual Speech
Date Added 2026-07-21
  • ✓16 languages from one hosted endpoint
  • ✓Flash tier hits roughly 300ms first-packet latency
  • ✓Plus tier trades speed for timbre fidelity
  • ✓Improved Chinese dialect coverage over the previous generation
  • ✓Streaming input and output via Model Studio
  • ⚠Hosted only — no downloadable weights, unlike the open Qwen3-TTS line
  • ⚠Requires an Alibaba Cloud Model Studio account
  • ⚠Alibaba has not published per-character pricing in the announcement
  • ⚠16 languages is narrower than the largest Western TTS catalogues
  • ⚠Flash and Plus differ enough that tier choice needs testing, not assumption

Call it through Alibaba Cloud Model Studio — there are no weights to download. Pick the Flash tier for live conversational work where first-packet latency governs the experience, and the Plus tier for narration or dubbing where listeners will notice timbre. Check current per-character pricing in Model Studio before committing to volume.

1 Choose the tier per workload, not per account — the API exposes both
2 Use Flash for anything a user waits on, Plus for anything they replay
3 Test your target dialect specifically, since dialect fidelity was the headline change
4 Stream output for long passages instead of waiting on full synthesis
5 Confirm pricing in Model Studio, as the tiers are billed differently

Live Multilingual Support Agent

Give a voice agent low-latency speech in a caller's own language.

STEPS:
  1. Detect the caller's language from the opening turn
  2. Route synthesis to the Flash tier for responsiveness
  3. Stream audio out as the reply is generated
  4. Fall back to text if latency budget is exceeded
  5. Log per-language quality for review

Localising a Video Library

Re-voice an existing catalogue across the supported languages.

STEPS:
  1. Export the source script per segment
  2. Synthesise with the Plus tier for fidelity
  3. Spot-check timbre and pacing per language
  4. Align audio to the original timings
  5. Publish with per-language audio tracks
Paid

Requires a paid subscription.

View Qwen-Audio-3.0-TTS Alternatives (2026) →

Compare Qwen-Audio-3.0-TTS with 5+ similar text → audio AI tools.

❓
Q

Is Qwen-Audio-3.0-TTS free?

A

No, Qwen-Audio-3.0-TTS requires a paid subscription.

Q

Does Qwen-Audio-3.0-TTS have an API?

A

Yes, Qwen-Audio-3.0-TTS offers an API for programmatic integration.

Q

What is Qwen-Audio-3.0-TTS best for?

A

Qwen-Audio-3.0-TTS is best for Multilingual Speech. Qwen-Audio-3. The two-tier split is the useful part: most TTS vendors make you pick latency or fidelity across the whole account, and this exposes both behind one API. The 300ms-level first-packet latency on the Flash tier is fast enough for live conversational use.

Q

What platforms does Qwen-Audio-3.0-TTS support?

A

Qwen-Audio-3.0-TTS supports api.

Q

Is Qwen-Audio-3.0-TTS open source?

A

No, Qwen-Audio-3.0-TTS is not open source.

Q

How do I get started with Qwen-Audio-3.0-TTS?

A

Call it through Alibaba Cloud Model Studio — there are no weights to download. Pick the Flash tier for live conversational work where first-packet latency governs the experience, and the Plus tier for narration or dubbing where listeners will notice timbre. Check current per-character pricing in Mod...

Q

How do I generate audio with Qwen-Audio-3.0-TTS?

A

Qwen-Audio-3.0-TTS creates audio from text descriptions. Enter prompts describing the type of audio you want (voice, music, sound effects) along with style, tone, and duration details.

🏷️

Work on Qwen-Audio-3.0-TTS? You're hand-reviewed in our directory. Add this badge to your site — it links back to this profile.

Featured on CuratedAI