Announcing SpaceXAI's newest speech-to-text model, with unparalleled accuracy and cost effectiveness.
Today we're releasing Grok Voice Transcribe 2.0, our latest speech-to-text model. Across our real-world evaluations, Grok Voice Transcribe 2.0 is one of the most accurate transcription models available today and twice as accurate as Grok Voice Transcribe 1.0, at the same price.
Grok Voice Transcribe 2.0 is built on the audio foundation model behind Grok Voice. Grok Voice already powers tens of thousands of customer-support calls a day, transcribes millions of hours of video narration, and runs voice agents in physical products, including the Grok assistant in Tesla vehicles. It is trained on a unique dataset of live, noisy, multilingual audio recorded across a diverse set of environments and refined with post-training.
The result is one of the most accurate transcription models for speech in real-world settings.
Most transcription models do well on clean, single-speaker audio. Real-world audio is harder: flaky phone lines, competing voices, local accents, and phone numbers or email addresses read aloud. We built Grok Voice Transcribe 2.0 for the hardest audio across conditions and environments.
On the public Artificial Analysis leaderboard, Grok Voice Transcribe 2.0 ranks first for accuracy among 32 streaming models.
Upper and to the right is better.
In addition to public benchmarks, we measure word error rate on four internal sets drawn from production traffic: telephony audio from customer-support calls, conversations with Grok, spoken credentials such as account codes and email addresses, and short multilingual voice commands. Grok Voice Transcribe 2.0 improves on Grok Voice Transcribe 1.0 across all four, and on telephony it leads every model we tested.
Customer support calls · English
Conversations with Grok · English
Phone numbers, emails, addresses · English
Voice-assistant utterances · 19 languages
Grok Voice Transcribe 2.0 transcribes dozens of languages, detects the language automatically, and follows mid-recording switches in a single pass. Multilingual accuracy is its largest improvement over Grok Voice Transcribe 1.0.
Word Error Rate (%). Lower is better.
Short phrases such as in-car commands give the model little context to identify the language. On our short-phrase set, word error rate drops from 20.6% to 6.8%.
Grok Voice Transcribe 2.0 supports advanced configuration and controls. Existing Speech-to-Text API integrations get the accuracy improvement with no code changes:
Atlassian Loom is widely used to record and share screen recordings. Atlassian found Grok Voice Transcribe 2.0 more accurate than its existing solution and now uses it to transcribe every video. Accurate transcripts open new AI workflows: record an action plan in Loom, pipe the transcript into Cursor, and it makes the code updates directly.
“We've always believed the best way to move work forward is to capture context once and let it flow everywhere. With Grok powering Loom's speech-to-text and Cursor turning that into code, we're closing the loop from context to code: record what you mean, and the work gets done. It's a glimpse of where AI-assisted development is headed.”
Grok Voice Transcribe 2.0 pricing is identical to Grok Voice Transcribe 1.0. Batch transcription remains $0.10 per hour of audio and streaming $0.20 per hour, with diarization, timestamps, and key terms included.
USD per hour of audio
USD per hour of audio
Grok Voice Transcribe 2.0 will soon be the default in the Speech-to-Text API, and Grok Voice Transcribe 1.0 will be deprecated in the coming weeks. To stay on it during the transition, pin grok-voice-transcribe-1.0.