Step1V
Step1V established StepFun’s product and research path from language models toward native visual understanding.
FAMILY / STEPFUN
StepFun’s general, multimodal, and agent model line from Step1V and Step2 through Step 3, Step 3.5 Flash, Step 3.7 Flash, and Step 5 Preview.
01 / RELEASE SEQUENCE
02 / GENERATIONS
Step1V established StepFun’s product and research path from language models toward native visual understanding.
Step-1.5V advanced image and video understanding from the early Step1V milestone into a distinct product model.
Step-2 moved StepFun’s scaled MoE foundation capability into a formal product milestone; specific parameter figures are not backfilled from secondary material here.
Step-R1-V-Mini shows StepFun’s early route for extending reinforcement-learning reasoning to visual tasks.
Step 3 is a key point where StepFun’s public-weight path moved from early product models to efficient multimodal reasoning and model-system co-design.
STEP3-VL-10B brings Step’s visual reasoning to a more deployable scale and uses parallel coordinated reasoning to scale test-time compute.
Step 3.5 Flash combines sparse activation, sliding-window attention, and multi-token prediction into an open model for real-world agent tasks.
Step 3.7 Flash extends Step 3.5 Flash’s high-throughput agent path to native vision, search, and longer-horizon tool execution.
Step 5 Preview advances StepFun’s model line to larger-scale long-horizon agents and professional knowledge work, while its open weights are not yet available.