Model index
StepFun·StepAudio

Step-Audio 2

An end-to-end audio-understanding and speech-conversation release group spanning audio semantics, paralinguistic information, tool use, and multimodal retrieval.

CurrentLanguageReasoningMultimodalOpen weightsRelease group
MODEL NUMBER343
First public date
2025-07-23
Organization
StepFun
Family
StepAudio
01

Specifications

Only published specifications are shown.

This record covers multiple tiers. Modalities and capabilities describe the group coverage; support and licensing depend on the individual tier.

Input modalities
Text · Audio
Output modalities
Text · Audio
Open weights
Yes
License
Apache-2.0
Status
Current
Step-Audio 2 mini · Apache-2.0 · in: Text/Audio · out: Text/AudioStep-Audio 2 mini Base · Apache-2.0 · in: Text/Audio · out: Text/AudioStep-Audio 2 mini Think · Apache-2.0 · in: Text/Audio · out: Text/Audio
02

Overview

Overview

An end-to-end audio-understanding and speech-conversation release group spanning audio semantics, paralinguistic information, tool use, and multimodal retrieval.

Why it mattered

Step-Audio 2 moved StepFun’s speech capability into reproducible open weights and inference code.

03

Architecture & capabilities

Model capabilities

Native audioReasoningTool use
04

Release history

Announcement

Step-Audio 2 technical report and audio-model line were introduced.

View source
Open weights

Step-Audio 2 mini and mini Base open weights were released.

View source
Major update

Step-Audio 2 mini Think was released.

View source
05

Related models

Select a card to view the related model.

No related models recorded yet

06

Sources

Suggest a correction

Copy this template, add the field, proposed value, and primary source, and share it with the archive maintainer.