Overview
Uses a Thinker-Talker architecture to understand text, images, audio, and video and stream text and speech.
Uses a Thinker-Talker architecture to understand text, images, audio, and video and stream text and speech.
Only published specifications are shown.
Uses a Thinker-Talker architecture to understand text, images, audio, and video and stream text and speech.
Qwen2.5-Omni-7B released model weights.
View sourceSelect a card to view the related model.
Copy this template, add the field, proposed value, and primary source, and share it with the archive maintainer.