Overview
An end-to-end audio-understanding and speech-conversation release group spanning audio semantics, paralinguistic information, tool use, and multimodal retrieval.
An end-to-end audio-understanding and speech-conversation release group spanning audio semantics, paralinguistic information, tool use, and multimodal retrieval.
Only published specifications are shown.
This record covers multiple tiers. Modalities and capabilities describe the group coverage; support and licensing depend on the individual tier.
An end-to-end audio-understanding and speech-conversation release group spanning audio semantics, paralinguistic information, tool use, and multimodal retrieval.
Step-Audio 2 moved StepFun’s speech capability into reproducible open weights and inference code.
Step-Audio 2 technical report and audio-model line were introduced.
View sourceStep-Audio 2 mini and mini Base open weights were released.
View sourceStep-Audio 2 mini Think was released.
View sourceSelect a card to view the related model.
No related models recorded yet
Copy this template, add the field, proposed value, and primary source, and share it with the archive maintainer.