MachinoAI Research
Grok 4.7 Benchmarks: Vendor Tables, Independent Runs, and Harness Effects
Grok 4.7 has different benchmark results depending on evaluator, harness and effort setting, so provenance must be preserved instead of collapsing scores into one leaderboard.
TL;DR
SpaceXAI reports 46.3% CursorBench 4.0, 71.0% DeepSWE v1.1 at high effort, 64.0% EEBench in the launch table, 1,657 AA Briefcase Elo, 37.6% Terminal-Bench 4.0, 19.6% Harvey Legal Agent Benchmark and 56.7% HealthBench Professional. Artificial Analysis independently reports 46 Intelligence Index, 1,657 AA-Briefcase, 1,695 GDPval-AA and 56 Coding Agent Index with Grok Build.
Why It Matters
Benchmark numbers for agentic systems encode the harness and effort budget. The same weights can produce materially different Terminal-Bench results in Grok Build versus a standardized evaluator.
Research Brief
The shortest useful explanation.
SpaceXAI reports 46.3% CursorBench 4.0, 71.0% DeepSWE v1.1 at high effort, 64.0% EEBench in the launch table, 1,657 AA Briefcase Elo, 37.6% Terminal-Bench 4.0, 19.6% Harvey Legal Agent Benchmark and 56.7% HealthBench Professional. Artificial Analysis independently reports 46 Intelligence Index, 1,657 AA-Briefcase, 1,695 GDPval-AA and 56 Coding Agent Index with Grok Build.
Core Explanation
A benchmark record preserving the official comparison matrix, model-card results, independent measurements, pricing, evaluation settings and harness differences for Grok 4.7.
Why It Matters
Benchmark numbers for agentic systems encode the harness and effort budget. The same weights can produce materially different Terminal-Bench results in Grok Build versus a standardized evaluator.

Grok 4.7 promotional visual
Direct image asset used as the article hero visual. This is an original third-party promotional image about Grok 4.7, not an official xAI benchmark figure.
Section 01
Problem
01. Problem
AI benchmark tables increasingly combine different model settings, evaluation harnesses, benchmark versions and provider tooling. Grok 4.7 is a concrete example: the launch table, the model card and Artificial Analysis expose different numbers for some of the same named benchmarks. A useful benchmark record therefore needs provenance before aggregation.
Section 02
How It Works
03. How It Works
SpaceXAI’s headline benchmark table compares Grok 4.7 xHigh with Grok 4.6 High, GPT-5.6 Sol Max and Fable 5.1 Max. The model card adds externally run evaluations and names the harness for several tests. Artificial Analysis then measures Grok in Grok Build for its Coding Agent Index and also runs separate standardized evaluations.
Section 03
Architecture
05. Architecture
Benchmark architecture is part of the measurement. Grok 4.7 can be evaluated in Grok Build while Artificial Analysis can use a standardized coding harness. DeepSWE has its own evaluator and CursorBench is tied to long-running coding sessions. The measured system is therefore model plus harness, tool permissions, retries and grading—not weights alone.
Section 04
Experiments
06. Experiments
The official comparison and independent results are preserved below. Official launch values come from SpaceXAI. Independent measurements come from Artificial Analysis. The official page also has interactive professional-work and cost-effectiveness visuals; their numeric values are preserved separately so they are not confused with the headline matrix.
References
Introducing Grok 4.7
Official launch announcement with training changes, benchmark comparison, pricing, availability and safety claims.
Grok 4.7 model card
Official benchmark and safety evaluation details, including additional model-card runs.
Benchmarking Grok 4.7
Independent measurements including Intelligence Index, Coding Agent Index, token use and task duration.
Official Grok 4.7 visual set
The official launch page contains the published launch artwork plus benchmark/professional-work and cost-effectiveness visuals. The verified page fetch did not expose stable direct image asset URLs, so no unverified figure URL was inserted.
SpaceXAI: Introducing Grok 4.7Grok 4.7 model documentationGrok 4.7 model cardArtificial Analysis: Benchmarking Grok 4.7