← In the News

Grok 4.7 gets a measurement its vendor did not run

Grok 4.7 Intelligence, Performance and Price Analysis · Artificial Analysis · September 2026

Machine-readable Download Markdown

Artificial Analysis has published its own evaluation of Grok 4.7, scoring the reasoning version at 46 on its Intelligence Index and ranking it 16th of the 655 models it tracks. Version 4.3.2 of that index bundles ten evaluations, including Terminal-Bench 4.0 for agentic coding and terminal use. The number practitioners should look at twice is not the rank. Running the index took Grok 4.7 240 million output tokens against a median of 92 million across comparable models, which Artificial Analysis reports as a verbosity rank of 140th out of 655. The model carries a 500k token context window.

Why it matters: Grok 4.7 arrived with every figure supplied by SpaceXAI. This is the first measurement of it we have seen that did not come from the vendor. It also moves the cost question off the price card: at about 2.6 times the median token spend per task, what a harness pays for this model is set by how much it writes, not by the per-token rate.