Overview
A vision-language release group with 3B, 7B, and 72B tiers for image, document, and video understanding.
A vision-language release group with 3B, 7B, and 72B tiers for image, document, and video understanding.
Only published specifications are shown.
This record covers multiple tiers. Modalities and capabilities describe the group coverage; support and licensing depend on the individual tier.
A vision-language release group with 3B, 7B, and 72B tiers for image, document, and video understanding.
Qwen2.5-VL released model weights.
View sourceSelect a card to view the related model.
Copy this template, add the field, proposed value, and primary source, and share it with the archive maintainer.