# xAI Releases Grok Voice Transcribe 2.0 With Major Accuracy Gains at Unchanged Pricing

> The September 18, 2026 update improves multilingual performance on noisy real-world audio while preserving batch and streaming options for enterprise developers.

*Published 2026-09-20 · By Marcus Vance*

Grok Voice Transcribe 2.0 is xAI's latest speech-to-text model that improves accuracy across dozens of languages by building on the audio foundation model used in Grok Voice.

xAI released Grok Voice Transcribe 2.0 on September 18, 2026. The model was added to the Speech-to-Text API alongside the existing grok-voice-transcribe-1.0 version. It supports both batch REST requests and real-time WebSocket streaming for files, URLs, or live audio.

The release focuses on real-world audio conditions. Training drew from live noisy multilingual production data that includes customer-support calls and Tesla vehicle recordings. This data source helps the model manage accents, background noise, and telephony distortions.

Version 2.0 incorporates automatic language detection. It also permits language switching during a single recording session. These features operate without separate passes or additional fees.

## What background explains the development of Grok Voice Transcribe 2.0?

The model extends prior work on Grok Voice audio capabilities. xAI trained the system on production-grade audio rather than clean studio recordings. This choice targets deployment scenarios where audio quality varies widely.

Customer-support telephony and in-vehicle environments provide representative training examples. Such data contains overlapping speech, variable signal quality, and rapid language shifts. The resulting model aims to maintain performance under these constraints.

The update arrives as demand grows for accurate voice input in development and enterprise tools. Integrations with platforms that convert speech into structured outputs benefit from lower error rates.

## What accuracy metrics does the new model report?

On an internal short-phrase multilingual evaluation covering 19 languages, the word error rate fell from 20.6 percent with version 1.0 to 6.8 percent with version 2.0. This improvement occurred on voice-assistant style utterances.

The model also achieved the top ranking among 32 streaming models on the Artificial Analysis public leaderboard. The leaderboard evaluates accuracy under standardized streaming conditions.

These results position the model as competitive with other frontier speech systems. The gains apply at the same price point as the prior version.

## What technical features come included with the release?

Diarization identifies distinct speakers within an audio stream. Word-level timestamps mark the start and end of each transcribed token. Key-term biasing allows users to emphasize specific vocabulary without extra configuration.

Both batch and streaming modes receive the full feature set. No separate charges apply for diarization, timestamps, or biasing. The API accepts either the new model identifier or the prior one.

- Batch REST transcription processes complete audio files or URLs in a single request.
- WebSocket streaming handles live audio input with incremental results.
- Automatic language detection occurs at the start of each session.
- Mid-recording language switching supports conversations that change languages.
- Speaker diarization and word timestamps are returned in the output JSON.

Key Metrics Comparison Between Grok Voice Transcribe VersionsMetricGrok Voice Transcribe 1.0Grok Voice Transcribe 2.0Internal Word Error Rate20.6%6.8%Leaderboard PositionNot ranked in top tier1 of 32 streaming modelsBatch Pricing$0.10 per hour$0.10 per hourStreaming Pricing$0.20 per hour$0.20 per hourIncluded FeaturesDiarization, timestamps, biasingDiarization, timestamps, biasing

## What market implications arise from the pricing and feature set?

Unchanged pricing removes a common barrier to adoption. Organizations can test or deploy the improved model without revising budgets. This structure supports incremental integration into existing pipelines.

Atlassian has incorporated the model into Loom workflows. The integration captures spoken context that later feeds into code generation tools. Such connections illustrate potential efficiency gains in software development cycles.

Enterprise users gain access to multilingual support without per-language surcharges. Automatic detection reduces the need for manual configuration in global teams. These elements align with broader demand for voice interfaces in productivity software.

## How have stakeholders responded to the announcement?

Atlassian executives highlighted the workflow benefits. The combination of improved transcription with downstream AI tools creates a direct path from spoken input to completed tasks.

> We've always believed the best way to move work forward is to capture context once and let it flow everywhere. With Grok powering Loom's speech-to-text and Cursor turning that into code, we're closing the loop from context to code: record what you mean, and the work gets done. It's a glimpse of where AI-assisted development is headed.Sanchan Saxena, SVP of Teamwork Collection, Atlassian

The reaction emphasizes practical integration rather than isolated benchmark scores. Organizations already using speech capture tools can expect measurable workflow compression.

## What steps follow the initial release?

Version 2.0 will become the default API setting in the near term. Users who require the prior model can pin their requests to grok-voice-transcribe-1.0. The older version faces deprecation within weeks.

Documentation updates at the official developer site detail the model identifiers and pricing tiers. Developers are advised to review migration paths before the default change takes effect.

Continued training on additional production data may yield further gains in specific domains. The current release establishes a baseline for subsequent iterations in the speech-to-text product line.

## Sources

1. [Today we're releasing Grok Voice Transcribe 2.0, our latest speech-to-text model. Across our real-world evaluations, Grok Voice Transcribe 2.0 is one of the most accurate transcription models available today and twice as accurate as Grok Voice Transcribe 1.0, at the same price.](https://x.ai/news/grok-voice-transcribe-2)
2. [Use `grok-voice-transcribe-1.0` or `grok-voice-transcribe-2.0`; the default is `grok-voice-transcribe-2.0`. ... REST pricing $0.10 / hr, Streaming pricing $0.20 / hr](https://docs.x.ai/developers/models/speech-to-text)
3. [xAI released Grok Voice Transcribe 2.0, an improved speech-to-text model with better accuracy across languages, priced at $0.10 per hour batch / $0.20 streaming, added alongside the prior version.](https://benchr.org/recent-releases)

---
Source: https://aiintelreport.com/frontier-models/grok-voice-transcribe-2-0-release
Index: https://aiintelreport.com/llms.txt · Full text: https://aiintelreport.com/llms-full.txt
