
Grok Voice Transcribe 2.0 was released on 18 September, priced at $0.10 per hour of audio for batch work and $0.20 per hour for streaming. Both figures are unchanged from the previous version.
The announcement is published under SpaceXAI, the name xAI has carried since July after SpaceX acquired it in February. Coverage still calls the company xAI, and its own footer does not.
Read the accuracy claim precisely
The company says the model is twice as accurate as Grok Voice Transcribe 1.0. That is a comparison with its own predecessor, not with anyone else’s model.
Against the field it makes a narrower claim, that it ranks first for accuracy among 32 streaming models on the public Artificial Analysis leaderboard, and describes itself as one of the most accurate rather than the most accurate. The hedge is the company’s own and is worth keeping.
Its internal charts compare against ElevenLabs Scribe v2 and Deepgram Nova-3. Those are vendor-run evaluations against vendor-chosen comparators, which is normal and is not the same as an independent test.
The price is the actual news
Holding price while claiming a doubling of accuracy is a statement about the market rather than the model. Speaker labelling, word-level timestamps and key term biasing are all included at no extra cost.
Those were billable features not long ago. Transcription is being priced as a commodity input, and the companies selling it are competing on what they can bundle.
The squeeze is general. OpenAI has pushed out new voice API models, and Microsoft’s in-house transcription model was benchmarked against Whisper, Gemini and ElevenLabs across 25 languages.
Where the advantage actually comes from
SpaceXAI is explicit about it. The model is trained on what it calls a unique dataset of live, noisy, multilingual audio recorded across a diverse set of environments.
It also describes the pipeline that produces that audio. Grok Voice powers tens of thousands of customer-support calls a day, transcribes millions of hours of video narration, and runs the assistant in Tesla vehicles.
Clean-speech accuracy is solved. What is left is telephony, accents, crosstalk and in-car commands, and the edge there belongs to whoever has the most real conversations flowing through their systems.
The evaluation sets are the detail to notice
The company says it measures word error rate on four internal sets drawn from production traffic. They are customer-support telephony, conversations with Grok, short multilingual voice commands, and spoken credentials.
That last set is described as account codes, phone numbers, email addresses and addresses read aloud. It is a standing evaluation corpus of people reading out their own identifying details.
Nothing in the announcement says how that material is obtained, retained or consented to. It may well be covered by enterprise terms, and the announcement does not say.
The European question that follows
A customer-support call has two parties and only one of them has a contract with the AI vendor. The caller reading out an account number has agreed to nothing.
Voice recordings are personal data in Europe, and a set built specifically around spoken credentials is about as sensitive as transcription data gets. Purpose limitation and retention are the obvious questions, and they are answerable.
The practice is not new, which is part of the point. People have been listening to supposedly automated transcripts for over a decade, and the disclosures have consistently lagged the practice.
Where the real gain is
The headline improvement is multilingual, and specifically short phrases. On the company’s own 19-language set of voice-assistant utterances, word error rate falls from 20.6% to 6.8%.
Short commands give a model almost no context to identify the language, which is exactly the in-car problem. The largest reported gain lands on the use case SpaceXAI owns outright through Tesla.
Multilingual is also where European vendors have concentrated. DeepL runs real-time voice translation across more than 40 languages, and ElevenLabs has listed transcription among five services on the UK government’s cloud framework.
The customer proof point
Atlassian says it found the model more accurate than its existing supplier and now transcribes every Loom video with it. That is a real switch by a named buyer, which is worth more than a benchmark chart.
Version 1.0 is being deprecated within weeks, with pinning available during the transition. Customers on the old model are being moved whether or not they asked.
What to watch
Watch whether anyone asks about the credentials corpus. It is the one genuinely novel disclosure in the announcement and nobody has picked it up.
Watch the price floor. At 10 cents an hour the model is nearly free, and the business is the audio flowing through it.