Model index
StepFun·Step

Step 3.7 Flash

A high-efficiency vision-language Flash model for real-world agents, with about 198B total parameters, about 11B active, 256K context, image understanding, and tool orchestration.

CurrentLanguageReasoningCodingAgentMultimodalOpen weights
MODEL NUMBER340
First public date
2026-05-29
Organization
StepFun
Family
Step
01

Specifications

Only published specifications are shown.

Total parameters
198B
Active parameters
11B
Architecture
MoE
Context
256,000 tokens
Input modalities
Text · Image
Output modalities
Text
Open weights
Yes
License
Apache-2.0
Status
Current
02

Overview

Overview

A high-efficiency vision-language Flash model for real-world agents, with about 198B total parameters, about 11B active, 256K context, image understanding, and tool orchestration.

Why it mattered

Step 3.7 Flash extends Step 3.5 Flash’s high-throughput agent path to native vision, search, and longer-horizon tool execution.

03

Specifications alongside predecessor

Compared with Step 3.5 Flash

PreviousStep 3.5 Flash196B · 256K ctx
CurrentStep 3.7 Flash198B · 256K ctx
04

Architecture & capabilities

Model capabilities

ReasoningCodingTool useVisionLong context

Connected tools

Web search

Unclassified tags

Fast action
05

Release history

Open weights

Step 3.7 Flash launched with open weights and deployment materials.

View source
06

Related models

Select a card to view the related model.

07

Sources

Suggest a correction

Copy this template, add the field, proposed value, and primary source, and share it with the archive maintainer.