Model index
SpaceXAI·Grok

Grok Voice Transcribe 2.0

SpaceXAI’s transcription model for batch and real-time speech recognition, improving multilingual accuracy, diarization, timestamps, and noisy real-world audio.

CurrentLanguageWeights not released

Grok Voice Transcribe 2.0 is available in the Speech-to-Text API and is now the default model; 1.0 remains available but will be deprecated in the coming weeks.

MODEL NUMBER121
First public date
2026-09-17
Organization
SpaceXAI
Family
Grok
01

Specifications

Only published specifications are shown.

Input modalities
Audio
Output modalities
Text
Open weights
No
Status
Current
02

Overview

Overview

SpaceXAI’s transcription model for batch and real-time speech recognition, improving multilingual accuracy, diarization, timestamps, and noisy real-world audio.

Why it mattered

It turns the audio foundation behind Grok Voice into a dedicated transcription API and replaces the first-generation model without requiring integration changes.

03

Specifications alongside predecessor

Compared with Grok Voice Transcribe 1.0

PreviousGrok Voice Transcribe 1.0 · ctx
CurrentGrok Voice Transcribe 2.0 · ctx
04

Architecture & capabilities

Model capabilities

Audio transcriptionReal-time audioSpeaker diarizationSpeech endpointingMultilingual
05

Release history

API launch

SpaceXAI made Grok Voice Transcribe 2.0 available in the Speech-to-Text API on September 17 and announced it would become the default transcription model; Grok Voice Transcribe 1.0 entered a deprecation transition. The launch article is dated September 18.

View source
06

Related models

Select a card to view the related model.

07

Sources

Suggest a correction

Copy this template, add the field, proposed value, and primary source, and share it with the archive maintainer.