# Grok 4.7 Ties MiMo-V2.6-Pro at Top of Cyber Defense Benchmark as Refusals Sideline Rivals

> A new composite benchmark for defensive AI security places xAI and SpaceXAI/Xiaomi models ahead of higher-profile rivals that decline most tasks on safety grounds.

*Published 2026-09-29 · By Diane Okafor*

Grok 4.7 (xhigh) is the xAI model that ties MiMo-V2.6-Pro for first place on the Artificial Analysis Cyber Index, a composite defensive-security benchmark that scores audit, vulnerability discovery, and memory-safety tasks on real software repositories.

The leaderboard, published by Artificial Analysis, puts Grok 4.7 (xhigh) and MiMo-V2.6-Pro at 56, ahead of GPT-6 Luna (max) at 53 and GLM-5.3-Flash at 50. The result is the first composite read from a benchmark built specifically to evaluate agentic cyber defense rather than general coding or reasoning ability.

Artificial Analysis launched the index with the Cyber Index Alliance, a group that includes Collinear AI, IBM, NVIDIA, and Vercel. The benchmark runs on real software repositories and is limited to defensive tasks; according to the launch article, it does not request exploit building. That design is intended to make the evaluation usable by security teams without raising the safety concerns associated with offensive benchmarks.

The practical effect is a ranking that can separate model capability from model willingness. Some of the strongest general-purpose models appear lower on the Cyber Index not because they failed the tasks, but because they refused them on safety grounds. Artificial Analysis reports that GPT-6 Sol (max), GPT-6 Astra (max), Claude Fable 5.1, Claude Opus 5.5, and Qwen3.8 variants refuse 98-100% of tasks, which lowers their composite scores.

## What does the Artificial Analysis Cyber Index measure?

The Cyber Index is a composite of three evaluations, each aimed at a different stage of defensive security work. CWE-Bench-AA covers audit and patch; DeepsecBench-AA covers vulnerability discovery; CyberGym-E2E-AA covers end-to-end memory-safety find, reproduce, and patch workflows. According to Artificial Analysis, the index measures only defensive cyber work on real software repositories.

*Artificial Analysis Cyber Index components (Source: Artificial Analysis)*

| Component | Focus | Task |
| --- | --- | --- |
| CWE-Bench-AA | Audit and patch | Identify and fix weaknesses in real code |
| DeepsecBench-AA | Vulnerability discovery | Find exploitable defects in repositories |
| CyberGym-E2E-AA | Memory safety | Find, reproduce, and patch memory-safety issues end to end |

Because the three components are distinct, the composite score is not a single measure of coding skill. A model could be strong at audit and patch but weaker at end-to-end reproduction. Enterprise teams that deploy security agents need component-level detail to match a model to a workflow, and Artificial Analysis says the index documentation includes component benchmarks alongside the composite.

## Which models lead the cyber defense leaderboard?

The published leaderboard shows Grok 4.7 (xhigh) and MiMo-V2.6-Pro tied at 56, with GPT-6 Luna (max) at 53 and GLM-5.3-Flash at 50. Developer Tech reported the same ordering, putting Grok 4.7 (xhigh) and MiMo-V2.6-Pro at the top with matching scores of 56 percent and GPT-6 Luna at 53 percent.

The top-two tie is notable because it places an xAI model alongside MiMo-V2.6-Pro, a model tied to SpaceXAI and Xiaomi. It also leaves GPT-6 Luna (max) in third despite GPT-6 models being positioned as frontier systems. The leaderboard shows that security-specific benchmarks can produce a different competitive order than general reasoning tests.

## Why do safety refusals suppress scores for frontier rivals?

Refusal rates are the clearest explanation for the gap between general capability and Cyber Index position. According to Artificial Analysis, GPT-6 Sol (max), GPT-6 Astra (max), Claude Fable 5.1, Claude Opus 5.5, and Qwen3.8 variants refuse 98-100% of tasks on safety grounds. Those refusals lower their scores even though the benchmark is defensive.

The refusal pattern is not uniform across all models. GPT-6 Luna (max) reached 53, which suggests that at least one GPT-6 configuration was able to complete a meaningful share of defensive tasks. The contrast within the same model family is a useful signal for enterprises: the specific model version and configuration, not just the lab, determines whether a security agent will be usable.

Artificial Analysis says GPT-6 models in particular refuse most CyberGym-E2E-AA tasks, which is the component that asks models to find, reproduce, and patch memory-safety issues. Because that component is central to the index, high refusal rates there have an outsized effect on the composite. A model that declines end-to-end tasks will not appear at the top of the leaderboard even if it scores well on audit-only components.

## What does the benchmark mean for enterprise security teams?

For enterprises, the Cyber Index addresses a recurring problem in AI procurement: general benchmarks do not tell a security team whether a model can handle real defensive work. The index's focus on audit, vulnerability discovery, and memory-safety repair maps more directly to security operations than coding or math leaderboards.

The composite score is useful for quick comparison, but procurement decisions should be based on the component breakdown. A team that runs patch review may prefer a model with strong CWE-Bench-AA performance, while a team that investigates reported vulnerabilities may weight DeepsecBench-AA more heavily. The refusal data is equally important, because a model that declines tasks will fail in production even if its theoretical capability is high.

Enterprises should also treat the benchmark as a starting point rather than a certification. The tasks run on real software repositories, but production environments have their own permissions, tools, and policy constraints. Security teams should run their own evaluations in a sandbox that mirrors their stack before deploying an AI agent.

- Review component scores, not just the composite, to match model strengths to specific security workflows.
- Test refusal behavior in a sandbox that mirrors production repositories and permissions.
- Confirm that benchmark tasks reflect defensive work and do not require exploit construction.
- Re-check leaderboard positions regularly, because the index is new and model versions are changing quickly.

## How are benchmark results changing the competitive picture?

The Cyber Index introduces a new axis of competition for AI labs. xAI's Grok 4.7 (xhigh) and MiMo-V2.6-Pro now hold the top defensive-security position, while several higher-profile frontier models are held back by refusal behavior. That changes the story that labs can tell enterprise buyers about security readiness.

The result also puts pressure on labs with high refusal rates to decide whether defensive security tasks should be treated differently from offensive ones in their safety policies. Artificial Analysis's benchmark does not request exploit building, which narrows the policy question: if a model refuses defensive work, the reason is a safety posture, not an inherent technical limit.

Independent coverage by Developer Tech confirmed the top-two result and framed the benchmark as a test of defensive AI security models. The agreement between the primary index and the trade press reduces the chance that the leaderboard ordering is a single-source artifact.

## What should enterprises watch next?

The Cyber Index is new, and the alliance behind it includes Collinear AI, IBM, NVIDIA, and Vercel. That institutional backing suggests the benchmark will be updated as models are released, and the methodology is likely to be refined as security teams report what the component scores do and do not predict.

Enterprises should watch for three things: component-level scores for new model versions, refusal rates on defensive tasks, and the cost of running top-scoring models at scale. Artificial Analysis says the index documentation includes methodology, safety blocks, costs, and component benchmarks, which gives buyers more than a single leaderboard number.

The next round of results will show whether Grok 4.7 (xhigh) and MiMo-V2.6-Pro can hold the top position as rivals tune their safety behavior or improve their defensive tooling. Until then, the practical takeaway for enterprises is to evaluate security agents on defensive tasks, measure refusal behavior, and map component scores to the specific workflows they need to automate.

## Sources

1. [Grok 4.7 (xhigh) and MiMo-V2.6-Pro tie for first on the Artificial Analysis Cyber Index with a score of 56.](https://artificialanalysis.ai/evaluations/artificial-analysis-cyber-index)
2. [The Cyber Index Alliance includes Collinear AI, IBM, NVIDIA, and Vercel. The index measures only defensive cyber work on real software repositories; no exploit building is requested. GPT-6 models refuse most CyberGym-E2E-AA tasks.](https://artificialanalysis.ai/articles/artificial-analysis-cyber-index)
3. [Grok 4.7 (xhigh) and MiMo-V2.6-Pro top the leaderboard with matching scores of 56 percent. GPT-6 Luna followed at 53 percent.](https://www.developer-tech.com/news/artificial-analysis-cyber-index-tests-defensive-ai-security-models/)

---
Source: https://aiintelreport.com/enterprise-ai/grok-4-7-ties-cyber-defense-lead
Index: https://aiintelreport.com/llms.txt · Full text: https://aiintelreport.com/llms-full.txt
