StepAudio 3
StepAudio 3 brings speech recognition, realtime conversation, and audio generation into one product generation.
FAMILY / STEPFUN
StepFun’s audio line spanning Step-Audio 2, audio reasoning, and StepAudio 3 for realtime interaction, recognition, and generation.
01 / RELEASE SEQUENCE
02 / GENERATIONS
StepAudio 3 brings speech recognition, realtime conversation, and audio generation into one product generation.
Step-Audio 2 moved StepFun’s speech capability into reproducible open weights and inference code.
Step-Audio-R1 extends StepFun’s reasoning line from text and vision to audio understanding.
The open-weight update to Step-Audio-R1, with inference code and weights released on 2026-01-14 in the official repository.
The audio-reasoning research update to Step-Audio-R1, with a technical report and accompanying evaluation benchmarks released on 2026-04-29.