Frontier Models
Grok 4.7 Delivers Self-Verifying Coding Gains at Unchanged Grok 4.6 Pricing
SpaceXAI's latest release enhances self-checking and long context capabilities for coding and knowledge tasks while preserving the token pricing from the previous model version.
Grok 4.7 is xAI's frontier model update from Grok 4.6 that uses a larger base model and extended reinforcement learning to improve performance on complex coding and knowledge tasks.
The release comes after some timeline adjustments but delivers notable advances in self-verification mechanisms that allow the model to check its own outputs more thoroughly during extended tasks. SpaceXAI trained the model with a longer reinforcement learning phase focused on difficult problems that require many hours of work. This training strategy contributes to higher scores on benchmarks designed for software engineering and terminal operations. The model also improves document and presentation creation capabilities while running at the same price with a faster variant available in select environments.
What background led to the development of Grok 4.7?
Grok 4.7 builds directly on the architecture of Grok 4.6 but incorporates a new larger base model to handle increased complexity. The company applied additional training resources to reinforcement learning stages that emphasize multi-hour tasks. This approach addresses limitations in previous versions regarding sustained performance on intricate coding projects. Developers have noted the need for models that can maintain coherence over long contexts and verify their reasoning steps independently.
The May 2026 knowledge cutoff ensures the model incorporates recent information up to that point. Availability in tools like Cursor positions it for immediate use in professional workflows. The emphasis on self-verification allows the model to identify potential errors in code generation before final output. This feature reduces debugging time for users working on complex software projects.
What technical specifications define Grok 4.7?
The model features a context window of 500,000 tokens which supports processing of extensive documents and codebases in a single session. Training data includes a harder mix of tasks compared to the prior iteration. SpaceXAI integrated a new safeguard stack that balances stronger refusals for inappropriate requests with low refusal rates for valid queries. Long-context handling receives particular attention to maintain accuracy across large inputs.
Improvements target self-verification during task execution, allowing the model to identify and correct errors in its generated code or analyses. Document and presentation creation capabilities also see enhancements for professional output generation. The new larger base model contributes to these gains through more robust pattern recognition in extended sequences.
Which benchmark scores highlight the performance improvements?
Performance metrics indicate clear gains in areas relevant to real-world coding applications. According to data from SpaceXAI, the model reaches 46.3 percent on CursorBench 4.0 for software engineering tasks. This represents an increase from 40.4 percent achieved by Grok 4.6 on the same benchmark. The results demonstrate the impact of the extended reinforcement learning run.
| Benchmark | Grok 4.6 | Grok 4.7 | Source Publisher |
|---|---|---|---|
| CursorBench 4.0 | 40.4% | 46.3% | SpaceXAI |
| Terminal-Bench 4.0 | 20.3% | 38.0% | SpaceXAI |
| EEBench | N/A | 64.0% | SpaceXAI |
On Terminal-Bench 4.0 which measures multi-hour terminal work the score nearly doubles to 38.0 percent from 20.3 percent. The electrical engineering benchmark EEBench shows a 64.0 percent score for the new model. These results underscore the effectiveness of the extended training regimen on challenging task distributions. Long context handling benefits from the larger window size during these evaluations.
How does pricing and availability compare to previous offerings?
Pricing remains unchanged from Grok 4.6 at two dollars per million input tokens and six dollars per million output tokens. A fast variant offers twice the output speed but comes at twice the price and is accessible in Cursor and Grok Build. This structure allows users to choose between standard and accelerated performance based on their needs. The model launched on September 21, 2026 and became available immediately in the listed platforms.
- Cursor
- Grok Build
- Grok API
- Third-party coding harnesses
- Model routers
- Cloud platforms
Integration with the Grok API facilitates custom applications and third party tools. Cloud platforms provide scalable deployment options for enterprise users. The unchanged pricing supports broader accessibility across developer communities and organizations seeking cost effective frontier model options.
What are the market and stakeholder implications of this release?
The decision to keep pricing the same while delivering performance gains strengthens the competitive position of xAI in the frontier model space. Software developers and knowledge workers gain access to enhanced tools without incurring higher costs. The focus on self-verification may reduce the need for extensive human oversight in coding projects. Stakeholders in the AI industry will monitor how these improvements influence adoption rates in agentic and coding applications.
The longer context window supports more complex workflows that involve large code repositories. Improved presentation creation features expand utility beyond pure coding into documentation and reporting tasks. Industry observers note that such features contribute to broader acceptance of AI tools in regulated sectors.
What expert reactions have emerged regarding Grok 4.7?
Elon Musk described the model as a strong combination of intelligence, speed and low cost in a public statement. The official announcement from SpaceXAI highlights its position as the most capable model for coding and knowledge work to date. The combination of performance gains and maintained pricing receives positive attention from the developer community.
Grok 4.7 is our most capable model for coding and knowledge work. It works longer on difficult tasks, checks its own work more carefully, and comes with our best-calibrated safeguards to date. Served at the same price and speed as Grok 4.6, it is highly competitive in its class.SpaceXAI
Users in coding environments like Cursor report seamless integration and noticeable improvements in task completion accuracy. The new safeguard stack with stronger refusals and jailbreak resistance maintains low refusal rates for legitimate uses.
What developments can be expected next from xAI in this area?
Future iterations may further refine the safeguard mechanisms and expand context capabilities beyond the current 500,000 tokens. Continued emphasis on reinforcement learning for long duration tasks could yield additional gains in agentic performance. The company may introduce more specialized variants for specific industry applications. Integration with additional platforms and tools is likely as adoption grows.
Monitoring of real world usage will inform subsequent training adjustments. The current release sets a foundation for models that require less intervention in complex problem solving scenarios. The balance of intelligence, speed and cost positions the model for sustained relevance in competitive landscapes.
Frequently asked
When was Grok 4.7 released by SpaceXAI?
Grok 4.7 was released on September 21, 2026.
What is the pricing for Grok 4.7 tokens?
The pricing is $2 per million input tokens and $6 per million output tokens, the same as Grok 4.6.
Which platforms support Grok 4.7 immediately?
It is available in Cursor, Grok Build, the Grok API, third-party coding harnesses, model routers, and cloud platforms.
What benchmark improvements are reported for Grok 4.7?
It scores 46.3% on CursorBench 4.0 and 38.0% on Terminal-Bench 4.0 according to SpaceXAI.
Sources
- SpaceXAI — Grok 4.7 scores 46.3% on CursorBench 4.0, up from 40.4% for Grok 4.6, and 38.0% on Terminal-Bench 4.0. The model is priced at $2 per million input tokens and $6 per million output tokens.
- Cursor — Grok 4.7 scores 46.3% on CursorBench 4.0, up from 40.4% for Grok 4.6, and nearly doubles Grok 4.6 on Terminal-Bench 4.0 (20.3% → 38.0%). It was trained with a longer reinforcement learning run on a harder mix of tasks.
- SpaceXAI — Grok 4.7 has a context window of 500,000 tokens, input price $2.00 per million tokens, output price $6.00 per million tokens, and a fast variant billed at twice the standard token rates.
- X — Grok 4.7 is a strong combination of intelligence, speed & low cost
- X — Grok 4.7 works longer on difficult tasks, checks its work more carefully, and comes with our strongest safeguards to date.