Overview
A compact multimodal open model in the Step 3 line at about 10B scale, focused on visual perception, complex reasoning, and human-centric alignment.
A compact multimodal open model in the Step 3 line at about 10B scale, focused on visual perception, complex reasoning, and human-centric alignment.
Only published specifications are shown.
A compact multimodal open model in the Step 3 line at about 10B scale, focused on visual perception, complex reasoning, and human-centric alignment.
STEP3-VL-10B brings Step’s visual reasoning to a more deployable scale and uses parallel coordinated reasoning to scale test-time compute.
STEP3-VL-10B open model suite released.
View sourceSelect a card to view the related model.
Copy this template, add the field, proposed value, and primary source, and share it with the archive maintainer.