Back to events
Model releaseHigh global significanceConfirmed confidence

Qwen open-sources Qwen3-Omni

Qwen releases an open-weight, native end-to-end omni-modal model with real-time text and natural-speech output.

Event details

Qwen3-Omni jointly handles text, image, audio and video input and uses a Thinker–Talker architecture for low-latency text or speech responses.

Why it matters

It combines open weights, native omni-modal understanding and real-time speech generation in one model, lowering barriers to end-to-end voice research and deployment.

82/100Global significance score. Regional effects are recorded only when the evidence supports a meaningful difference.

Access notes

Weights, code and an online API are available. Speech input supports 19 languages and speech output supports 10.