xAI launches Grok Voice Transcribe 2.0, doubling transcription accuracy at the same price
xAI has released Grok Voice Transcribe 2.0, a new speech-to-text model the company says is twice as accurate as its predecessor at unchanged pricing. The model is available now through xAI's API for both batch and streaming transcription workloads.
What's new
Grok Voice Transcribe 2.0 adds several capabilities on top of the original Transcribe model:
- Speaker diarization at no additional cost, identifying which speaker said what in multi-person audio.
- Multichannel transcription across up to 8 audio channels in a single request.
- Key term biasing, letting developers supply up to 100 domain-specific terms (product names, jargon, acronyms) to improve recognition accuracy.
- Word-level timestamps and confidence scores for each transcribed word.
- Automatic text formatting for numbers, dates, currencies, phone numbers, and emails.
- Automatic language detection across dozens of supported languages.
Pricing is unchanged from the prior generation: $0.10 per hour of audio for batch transcription and $0.20 per hour for streaming. xAI says the accuracy gain comes at "the same price" as Transcribe 1.0.
In its announcement, xAI states: "Across our real-world evaluations, Grok Voice Transcribe 2.0 is one of the most accurate transcription models available today and twice as accurate as Grok Voice Transcribe 1.0, at the same price."
Context
The release is part of a broader build-out of xAI's Grok Voice product line, which also includes Grok Voice Think Fast, a separate speech-to-speech conversational model. Transcribe and Think Fast serve different jobs: Transcribe is a pure speech-to-text tool aimed at developers who need accurate transcripts (captioning, meeting notes, call-center analytics), while Think Fast handles real-time spoken conversation.
Speech-to-text has become a competitive front among frontier AI labs as voice interfaces spread into consumer and enterprise products. Rivals including OpenAI, Google, and dedicated speech vendors have all shipped competing transcription models with similar features — diarization, timestamps, and domain biasing are increasingly table stakes rather than differentiators.
Why it matters
A 2x accuracy jump at flat pricing is a meaningful improvement for any developer already budgeting for transcription at scale — call centers, video platforms, and note-taking apps are the most obvious beneficiaries, since word-error-rate reductions compound directly into lower manual-correction costs. The addition of free-tier diarization and multichannel support also narrows the feature gap with specialized transcription vendors, making Grok Voice Transcribe a more complete substitute for point solutions rather than a bare-bones API. For xAI, the release signals continued investment in the audio/voice stack alongside its chat models, reinforcing Grok Voice as a multi-product line rather than a single speech-to-speech offering.
Corroborating sources
- X
https://x.ai/news/grok-voice-transcribe-2
“Across our real-world evaluations, Grok Voice Transcribe 2.0 is one of the most accurate transcription models available today and twice as accurate as Grok Voice Transcribe 1.0, at the same price.”