Overview
A sparse MoE model for fast reasoning, coding, and agent workflows, with 196B total parameters, 11B active per token, and 256K context.
A sparse MoE model for fast reasoning, coding, and agent workflows, with 196B total parameters, 11B active per token, and 256K context.
Only published specifications are shown.
A sparse MoE model for fast reasoning, coding, and agent workflows, with 196B total parameters, 11B active per token, and 256K context.
Step 3.5 Flash combines sparse activation, sliding-window attention, and multi-token prediction into an open model for real-world agent tasks.
Step 3.5 Flash launched with open weights.
View sourceSelect a card to view the related model.
Copy this template, add the field, proposed value, and primary source, and share it with the archive maintainer.