Overview
An audio-native reasoning model using Modality-Grounded Reasoning Distillation so test-time compute scaling operates on acoustic information.
An audio-native reasoning model using Modality-Grounded Reasoning Distillation so test-time compute scaling operates on acoustic information.
Only published specifications are shown.
An audio-native reasoning model using Modality-Grounded Reasoning Distillation so test-time compute scaling operates on acoustic information.
Step-Audio-R1 extends StepFun’s reasoning line from text and vision to audio understanding.
Step-Audio-R1 inference code and weights were released.
View sourceSelect a card to view the related model.
Copy this template, add the field, proposed value, and primary source, and share it with the archive maintainer.