Research

MachinoAI Research

IMPORTANTAI Evaluation / Benchmarks

Grok 4.7 Benchmarks: Vendor Tables, Independent Runs, and Harness Effects

Grok 4.7 has different benchmark results depending on evaluator, harness and effort setting, so provenance must be preserved instead of collapsing scores into one leaderboard.

SpaceXAI, Artificial AnalysisSep 21, 2026SpaceXAI release + independent evaluation14 min read
AI Evaluation / BenchmarksGrok 4.7benchmarksTerminal-BenchCursorBenchDeepSWEAA-BriefcaseGDPvalEEBenchArtificial Analysis

TL;DR

SpaceXAI reports 46.3% CursorBench 4.0, 71.0% DeepSWE v1.1 at high effort, 64.0% EEBench in the launch table, 1,657 AA Briefcase Elo, 37.6% Terminal-Bench 4.0, 19.6% Harvey Legal Agent Benchmark and 56.7% HealthBench Professional. Artificial Analysis independently reports 46 Intelligence Index, 1,657 AA-Briefcase, 1,695 GDPval-AA and 56 Coding Agent Index with Grok Build.

Why It Matters

Benchmark numbers for agentic systems encode the harness and effort budget. The same weights can produce materially different Terminal-Bench results in Grok Build versus a standardized evaluator.

Research Brief

The shortest useful explanation.

SpaceXAI reports 46.3% CursorBench 4.0, 71.0% DeepSWE v1.1 at high effort, 64.0% EEBench in the launch table, 1,657 AA Briefcase Elo, 37.6% Terminal-Bench 4.0, 19.6% Harvey Legal Agent Benchmark and 56.7% HealthBench Professional. Artificial Analysis independently reports 46 Intelligence Index, 1,657 AA-Briefcase, 1,695 GDPval-AA and 56 Coding Agent Index with Grok Build.

vendor benchmarksindependent evaluationharness sensitivityreasoning efforttask costbenchmark provenanceconfidence intervalsGrok 4.7

Core Explanation

A benchmark record preserving the official comparison matrix, model-card results, independent measurements, pricing, evaluation settings and harness differences for Grok 4.7.

Why It Matters

Benchmark numbers for agentic systems encode the harness and effort budget. The same weights can produce materially different Terminal-Bench results in Grok Build versus a standardized evaluator.

Grok 4.7 promotional visual.

Grok 4.7 promotional visual

Direct image asset used as the article hero visual. This is an original third-party promotional image about Grok 4.7, not an official xAI benchmark figure.

Section 01

Problem

01. Problem

AI benchmark tables increasingly combine different model settings, evaluation harnesses, benchmark versions and provider tooling. Grok 4.7 is a concrete example: the launch table, the model card and Artificial Analysis expose different numbers for some of the same named benchmarks. A useful benchmark record therefore needs provenance before aggregation.

1Vendor result versus independent result.2xHigh versus high versus max effort.3Native harness versus standardized harness.4Benchmark-version changes can alter comparability.

Section 02

How It Works

03. How It Works

SpaceXAI’s headline benchmark table compares Grok 4.7 xHigh with Grok 4.6 High, GPT-5.6 Sol Max and Fable 5.1 Max. The model card adds externally run evaluations and names the harness for several tests. Artificial Analysis then measures Grok in Grok Build for its Coding Agent Index and also runs separate standardized evaluations.

1Launch table: one vendor-run comparison view.2Model card: benchmark rows plus runner/harness detail.3Artificial Analysis: independent measurements using its own methodology.

Section 03

Architecture

05. Architecture

Benchmark architecture is part of the measurement. Grok 4.7 can be evaluated in Grok Build while Artificial Analysis can use a standardized coding harness. DeepSWE has its own evaluator and CursorBench is tied to long-running coding sessions. The measured system is therefore model plus harness, tool permissions, retries and grading—not weights alone.

1Grok Build.2Standardized agent harnesses.3Benchmark-specific graders.4Effort and retry policies.

Section 04

Experiments

06. Experiments

The official comparison and independent results are preserved below. Official launch values come from SpaceXAI. Independent measurements come from Artificial Analysis. The official page also has interactive professional-work and cost-effectiveness visuals; their numeric values are preserved separately so they are not confused with the headline matrix.

1Official CursorBench 4.0: 46.3% vs Grok 4.6 40.4%, GPT-5.6 Sol 41.7%, Fable 5.1 51.8%.2Official DeepSWE v1.1: 71.0% at high effort vs 65.2%, 72.7%, 70.0%.3Official EEBench: 64.0% vs 53.0%, 39.4%, 56.4%; model card separately reports 66.0% xHigh for Grok 4.7.4Official AA Briefcase v1.1: 1,657 Elo vs 1,546, 1,487, 1,678.5Official Terminal-Bench 4.0: 37.6% on the live launch page; model card reports 38.0% xHigh.6Official Harvey Legal Agent Benchmark: 19.6% vs 15.8%, 2.5%, 6.7%.7Official HealthBench Professional: 56.7% vs 48.5%, 60.5%, 62.1%.8Official interactive GDPval: Grok 4.7 xhigh 1,695 Elo; Grok 4.6 high 1,605; Fable 5.1 max 1,735; GPT-6 Astra max 1,542.9Independent Artificial Analysis: Intelligence Index 46; Coding Agent Index 56 with Grok Build; AA-Briefcase 1,657; GDPval-AA 1,695; Terminal-Bench 4.0 33%; DeepSWE 73%; SWE-Atlas-QnA 63%.

References