Model index
StepFun·Step

STEP3-VL-10B

A compact multimodal open model in the Step 3 line at about 10B scale, focused on visual perception, complex reasoning, and human-centric alignment.

CurrentLanguageReasoningMultimodalOpen weights
MODEL NUMBER338
First public date
2026-01-14
Organization
StepFun
Family
Step
01

Specifications

Only published specifications are shown.

Total parameters
10B
Architecture
Undisclosed
Input modalities
Text · Image
Output modalities
Text
Open weights
Yes
License
Apache-2.0
Status
Current
02

Overview

Overview

A compact multimodal open model in the Step 3 line at about 10B scale, focused on visual perception, complex reasoning, and human-centric alignment.

Why it mattered

STEP3-VL-10B brings Step’s visual reasoning to a more deployable scale and uses parallel coordinated reasoning to scale test-time compute.

03

Architecture & capabilities

Model capabilities

ReasoningVision
04

Release history

Open weights

STEP3-VL-10B open model suite released.

View source
05

Related models

Select a card to view the related model.

06

Sources

Suggest a correction

Copy this template, add the field, proposed value, and primary source, and share it with the archive maintainer.