X moved video captions off self-hosted Whisper Large V3 onto Grok Speech to Text for multilingual performance, latency, and word error rate.
X is the global town square. With hundreds of millions of monthly active users, the platform operates at global scale.
Video is a large part of that. Videos that pass a minimum view threshold are automatically transcribed for closed captions — about 310,000 a day. Those captions matter for deaf and hard-of-hearing users, and for anyone browsing with the video muted.
X used to run transcription on a self-hosted Whisper Large V3. It fell short on latency, word error rate, and multilingual performance. Grok Speech to Text outperformed it on all three.
Speed and accuracy were not the only things X was looking for. Grok also generates high-precision word-level timestamps, so transcript chunks line up with spoken dialogue and are easier to follow on screen.
The X team fully migrated to Grok Speech to Text, so captions on the town square are faster, more accurate, and easier to read.