Model index
StepFun·StepAudio

Step-Audio-R1

An audio-native reasoning model using Modality-Grounded Reasoning Distillation so test-time compute scaling operates on acoustic information.

CurrentLanguageReasoningMultimodalOpen weights
MODEL NUMBER344
First public date
2025-11-27
Organization
StepFun
Family
StepAudio
01

Specifications

Only published specifications are shown.

Input modalities
Text · Audio
Output modalities
Text
Open weights
Yes
License
Apache-2.0
Status
Current
02

Overview

Overview

An audio-native reasoning model using Modality-Grounded Reasoning Distillation so test-time compute scaling operates on acoustic information.

Why it mattered

Step-Audio-R1 extends StepFun’s reasoning line from text and vision to audio understanding.

03

Architecture & capabilities

Model capabilities

Native audioReasoning
04

Release history

Open weights

Step-Audio-R1 inference code and weights were released.

View source
05

Related models

Select a card to view the related model.

06

Sources

Suggest a correction

Copy this template, add the field, proposed value, and primary source, and share it with the archive maintainer.