All families

FAMILY / STEPFUN

Step

StepFun’s general, multimodal, and agent model line from Step1V and Step2 through Step 3, Step 3.5 Flash, Step 3.7 Flash, and Step 5 Preview.

First model
2023
Models
9
Latest dated version
Step 5 Preview

01 / RELEASE SEQUENCE

Family release sequence

Compare versions →
  1. 2023-12Step1VLanguage · Multimodal
  2. 2024-07-04Step-1.5VLanguage · Multimodal
  3. 2024-07-04Step-2Language · Reasoning
  4. 2025-04-08Step-R1-V-MiniReasoning · Multimodal
  5. 2025-07-31Step 3Language · Reasoning
  6. 2026-01-14STEP3-VL-10BLanguage · Reasoning
  7. 2026-02-02Step 3.5 FlashLanguage · Reasoning
  8. 2026-05-29Step 3.7 FlashLanguage · Reasoning
  9. 2026-09-20Step 5 PreviewLanguage · Reasoning
9 versions
What changed ↓

02 / GENERATIONS

What changed

01

Step1V

Step1V established StepFun’s product and research path from language models toward native visual understanding.

02

Step-1.5V

Step-1.5V advanced image and video understanding from the early Step1V milestone into a distinct product model.

03

Step-2

Step-2 moved StepFun’s scaled MoE foundation capability into a formal product milestone; specific parameter figures are not backfilled from secondary material here.

04

Step-R1-V-Mini

Step-R1-V-Mini shows StepFun’s early route for extending reinforcement-learning reasoning to visual tasks.

05

Step 3

Step 3 is a key point where StepFun’s public-weight path moved from early product models to efficient multimodal reasoning and model-system co-design.

06

STEP3-VL-10B

STEP3-VL-10B brings Step’s visual reasoning to a more deployable scale and uses parallel coordinated reasoning to scale test-time compute.

07

Step 3.5 Flash

Step 3.5 Flash combines sparse activation, sliding-window attention, and multi-token prediction into an open model for real-world agent tasks.

08

Step 3.7 Flash

Step 3.7 Flash extends Step 3.5 Flash’s high-throughput agent path to native vision, search, and longer-horizon tool execution.

09

Step 5 Preview

Step 5 Preview advances StepFun’s model line to larger-scale long-horizon agents and professional knowledge work, while its open weights are not yet available.