What is the IQ of Grok-4.20?
Grok-4.20 scores 145 on the Mensa Norway IQ test — tied for #1 globally. We break down xAI's flagship model and its exceptional visual reasoning.
Grok-4.20: xAI's flagship model
Grok-4.20 is xAI's most advanced model, released in 2026. On the Mensa Norway IQ test (TrackingAI benchmark), Grok-4.20 Expert Mode scored 145 — tied with GPT-5.4 Pro Vision for #1 globally.
This is the highest IQ score ever recorded by any AI model on a standardized human IQ test, placing Grok-4.20 in the "genius" range — the top 0.1% of human intelligence.
- Expert Mode: 145 (perfect 36/36 on visual patterns)
- Standard Mode: ~135
- AI IQ composite (estimated): ~140
- Cost: 3.00/15.00 per 1M tokens
1. Benchmark breakdown
Mensa Norway IQ test
Grok-4.20 Expert Mode scored 145 on TrackingAI's Mensa Norway test — the maximum score, equivalent to a perfect 36/36. This ties it with GPT-5.4 Pro Vision for the highest IQ ever recorded by an AI.
In standard mode (without expert reasoning), Grok-4.20 scores approximately ~135 — still in the "very superior" range.
AI IQ composite
The AI IQ project doesn't yet have a full composite score for Grok-4.20, but based on its Mensa Norway performance and strong results on ARC-AGI and GPQA, its estimated composite IQ is ~140.
Visual pattern recognition
Grok-4.20 excels at visual-spatial reasoning. On the Mensa Norway test, which is heavily visual, it outperforms Claude Opus 4.8 (~130) and matches GPT-5.4 Pro Vision (145).
iqscore.io test
Grok-4.20 is estimated to score ~140-145 on the iqscore.io 36-question test, based on its performance on similar visual and logical reasoning tasks.
2. Where Grok-4.20 excels
Visual-spatial reasoning
#1 globally on Mensa Norway (145) — Grok-4.20 is unmatched at visual pattern recognition, spatial rotation, and abstract figure analysis.
Expert Mode reasoning
Grok-4.20's Expert Mode engages deeper reasoning chains, similar to OpenAI's o-series models. This boosts its IQ from ~135 (standard) to 145 (expert).
Real-time knowledge
Grok-4.20 has access to X (formerly Twitter) data in real-time, giving it an edge on current events and trending topics — though this doesn't directly affect IQ scores.
Mathematical reasoning
Strong performance on AIME and FrontierMath — Grok-4.20's mathematical reasoning is comparable to GPT-5.5 and Claude Opus 4.8.
Cost-effectiveness
At 3.00/15.00 per 1M tokens, Grok-4.20 is cheaper than Claude Opus 4.8 (5/25) and GPT-5.5 (5/30), while matching or exceeding their IQ on visual tests.
3. Where it falls short
Agentic tasks
Grok-4.20 is not yet ranked on GDPval-AA v2. Its agentic capabilities — multi-step real-world tasks — are expected to be strong but unproven at the level of Claude Opus 4.8 (1,890 Elo).
Hallucination rate
Grok-4.20 has a moderate hallucination rate — higher than Claude Opus 4.8 (35.9%) but lower than some older models. Its real-time data access can sometimes introduce factual errors.
Intelligence Index
Grok-4.20 is not yet fully ranked on the Artificial Analysis Intelligence Index v4.1. Its composite reasoning score is expected to be competitive but may trail Claude Opus 4.8 (55.7) and GPT-5.5 (54.8).
Emotional intelligence
Grok-4.20 is known for its "unfiltered" personality, which can be witty but less emotionally calibrated than Claude Opus 4.8 (~132 EQ).
4. Grok-4.20 vs other models
| Model | Mensa Norway IQ | AI IQ est. | Cost/1M tokens | Hallucination |
|---|---|---|---|---|
| Grok-4.20 | 145 | ~140 | 3/15 | moderate |
| GPT-5.4 Pro Vision | 145 | ~136 | — | moderate |
| Gemini 3.1 Pro | 141 | ~131 | 2/12 | low |
| Claude Opus 4.8 | ~130 | ~132-145 | 5/25 | lowest |
| GPT-5.5 | ~145 | ~136 | 5/30 | moderate |
| DeepSeek V4 Pro | ~111 | ~111 | 0.44/0.87 | moderate |
5. What does this mean for humans?
Grok-4.20 operates at an IQ of 145 in Expert Mode — higher than 99.9% of humans. But:
- IQ 145 = genius threshold: top 0.1%, only 1 in 1,000 people
- Visual genius ≠ general genius: excels at patterns, not necessarily at all cognitive tasks
- Expert Mode matters: the 10-point boost from standard to expert shows reasoning depth
- Real-time data is a double-edged sword: current but potentially less verified
- AI IQ is narrow: pattern recognition is one dimension of intelligence
Conclusion
- Genius tier: top 0.1% of human intelligence (1 in 1,000).
- Expert Mode boosts IQ by ~10 points: from ~135 (standard) to 145 (expert).
- Unmatched visual-spatial reasoning: #1 on visual pattern recognition.
- Cost-effective genius: 3/15 per 1M tokens — cheaper than Claude and GPT.
- Strong mathematical reasoning: comparable to GPT-5.5 and Claude.
- Not yet ranked on Intelligence Index — agentic capabilities unproven.
- Moderate hallucination rate — higher than Claude Opus 4.8.