Overview
The real-time audiovisual translation model in the Qwen3.8 family, using a Hybrid-MoE Thinker–Talker design for speaker separation, synchronized bilingual output, and contextual disambiguation.
The real-time audiovisual translation model in the Qwen3.8 family, using a Hybrid-MoE Thinker–Talker design for speaker separation, synchronized bilingual output, and contextual disambiguation.
QwenCloud provides `qwen3.8-livetranslate-flash-realtime` through the WebSocket Realtime API; it supports audio/image input, text/audio output, 60 source languages, and 29 audio-output languages.
Only published specifications are shown.
The real-time audiovisual translation model in the Qwen3.8 family, using a Hybrid-MoE Thinker–Talker design for speaker separation, synchronized bilingual output, and contextual disambiguation.
It extends Qwen3.8 from general omni-modal and agent workflows into a dedicated low-latency speech-translation branch.
QwenCloud released qwen3.8-livetranslate-flash-realtime with audio/image input, text/audio output, 60 source languages, 29 audio-output languages, and WebSocket Realtime API access.
View sourceSelect a card to view the related model.
Copy this template, add the field, proposed value, and primary source, and share it with the archive maintainer.