What is the IQ of Grok-4.20?
Grok-4.20 scores 145 on the Mensa Norway IQ test — tied for #1 among all AI models. We break down every benchmark and explain what makes xAI's model unique.
Grok-4.20: xAI's flagship model
Grok-4.20 is xAI's most advanced model, released in mid-2026. On the Mensa Norway IQ test, Grok-4.20 scores 145 — tied for #1 among all AI models, sharing the top spot with GPT-5.5. This places it in the "genius" range, smarter than 99.9% of humans. On the Artificial Analysis Intelligence Index, Grok-4.20's exact score is still being finalized, but it ranks in the top 5 globally.
What makes Grok-4.20 unique is its combination of raw intelligence and personality. Unlike Claude, GPT-5.5, and Gemini — which are designed to be helpful, harmless, and neutral — Grok is designed to be funny, irreverent, and willing to answer any question. This makes it the most distinctive AI model on the market: one that matches the smartest models on raw IQ while having a personality that no other model dares to have.
Key scores:
- Mensa Norway IQ (TrackingAI): 145 (#1 tied with GPT-5.5)
- AI IQ composite (VexoWire): ~140 estimated (#1 among AI models)
- Artificial Analysis Intelligence Index: top 5 (score being finalized)
- GDPval-AA v2: top 5
- Humanity's Last Exam: top 5 among AI models
- Hallucination rate: moderate
- Cost: 3.00/15.00 per 1M tokens
- Personality: unique — irreverent, humorous, uncensored
1. The raw IQ champion
Grok-4.20 scores 145 on the Mensa Norway IQ test — the highest possible practical score, tied with GPT-5.5. This is the raw intelligence benchmark: pure visual pattern recognition and abstract reasoning, the kind of fluid intelligence that IQ tests were designed to measure.
At IQ 145, Grok-4.20 is smarter than 99.9% of humans. If it were a person, it would be in the top 0.1% of human intelligence — the realm of Nobel laureates, Fields medalists, and once-in-a-generation thinkers. Only about 1 in 1,000 people achieves this score.
What's remarkable is that Grok-4.20 achieves this with a fundamentally different design philosophy than GPT-5.5. While GPT-5.5 was optimized for maximum performance on every benchmark, Grok-4.20 was designed to be a more "authentic" AI — one that speaks its mind, makes jokes, and doesn't shy away from controversial topics. The fact that it matches GPT-5.5 on raw IQ despite this different focus is a testament to xAI's engineering.
On the VexoWire AI IQ composite — which blends multiple IQ tests (Mensa Norway, Mensa Denmark, Raven's, WAIS-IV) — Grok-4.20's estimated composite IQ is ~140, making it the #1 AI model on this metric. This is higher than GPT-5.5 (~136) and significantly higher than Claude (~132) and Gemini (~131).
2. Benchmark breakdown
Mensa Norway IQ test: 145 (#1 tied)
The Mensa Norway IQ test consists of 36 visual pattern recognition questions. Grok-4.20 scores 145 — the practical maximum, equivalent to a perfect or near-perfect 36/36. This ties it with GPT-5.5 as the smartest AI model on a pure IQ test.
Grok-4.20's strong performance on this test reflects xAI's investment in visual processing and abstract reasoning. The Grok architecture was designed with strong multimodal capabilities, and its performance on visual pattern recognition tasks is among the best in the industry.
AI IQ composite (VexoWire): ~140 (#1)
The VexoWire composite IQ blends multiple tests (Mensa Norway, Mensa Denmark, Raven's, WAIS-IV). Grok-4.20's estimated composite IQ is ~140 — the highest among all AI models. This is higher than its Mensa Norway score alone would suggest, because Grok-4.20 also performs strongly on the verbal and analytical subtests included in the composite.
This makes Grok-4.20 the smartest AI model by composite IQ — a remarkable achievement for a model that's also known for its humor and irreverence. It proves that being funny doesn't mean being less intelligent.
Artificial Analysis Intelligence Index: top 5
On the Intelligence Index — the gold standard for comparing AI models — Grok-4.20 ranks in the top 5. Its exact composite score is still being finalized as the index is updated to include the latest model versions. However, it's expected to score competitively with Claude Opus 4.8 (55.7) and GPT-5.5 (54.8), though likely not surpassing them on agentic tasks.
GDPval-AA v2: top 5
On the agentic benchmark — which measures performance on complex, multi-step real-world tasks — Grok-4.20 ranks in the top 5. It's less capable than Claude Opus 4.8 (1,890 Elo) and GPT-5.5 (1,850 Elo) on sustained multi-step reasoning, but still highly competent. For most real-world tasks, Grok-4.20 is more than capable.
Humanity's Last Exam (HLE): top 5
On the hardest academic questions across all disciplines, Grok-4.20 ranks in the top 5 among AI models. This reflects xAI's strength in training on diverse, high-quality data — Grok has deep knowledge across science, mathematics, history, and culture.
3. What does an IQ of 145 mean?
To put Grok-4.20's IQ of 145 in context:
- 85-115 (68% of people): Average intelligence
- 115-130 (14%): Above average to high average
- 130-145 (2%): Gifted — qualifies for Mensa (top 2%)
- 145-160 (0.1%): Highly gifted / genius
- 160+ (0.003%): Profoundly gifted — Einstein, Hawking territory
Grok-4.20, at 145, sits at the genius threshold — smarter than 99.9% of humans. If it were a person, it would be in elite company: the top 0.1% of human intelligence, the realm of Nobel prizes and Fields medals. Only about 1 in 1,000 people achieves this score.
This is the same IQ as GPT-5.5, and significantly higher than Claude (~130) and Gemini (141). On a pure IQ test, Grok-4.20 and GPT-5.5 would tie for first place among AI models, with Gemini third and Claude fourth.
But as we've seen with the other models, IQ is only one dimension of intelligence. Grok-4.20's genius-level IQ is complemented by its unique personality — making it not just the smartest, but also the most distinctive AI model available.
4. Where Grok-4.20 excels
Raw IQ and visual reasoning
Grok-4.20 is tied for #1 on the Mensa Norway IQ test (145) and #1 on the VexoWire AI IQ composite (~140). For pure cognitive ability — pattern recognition, abstract reasoning, logical deduction — Grok-4.20 is the smartest AI model available. For visual reasoning tasks, it's matched only by GPT-5.5.
Personality and humor
This is Grok-4.20's defining feature. Unlike other frontier models that are carefully neutral and cautious, Grok is designed to be funny, irreverent, and authentic. It makes jokes, offers opinions, and doesn't shy away from controversial topics. For users who find Claude and GPT-5.5 too bland or cautious, Grok-4.20 is a breath of fresh air — an AI that feels like it has a personality, not just a function.
Uncensored responses
Grok-4.20 is designed to answer questions that other models refuse. While Claude, GPT-5.5, and Gemini have extensive safety filters that prevent them from discussing certain topics, Grok-4.20 is willing to engage with a much wider range of questions. This makes it valuable for research, journalism, and any application where information access matters more than caution.
Real-time information
Grok-4.20 has access to real-time information through X (formerly Twitter). This gives it an advantage for questions about current events, breaking news, and trending topics. While other models have knowledge cutoffs, Grok-4.20 can access the latest information and provide up-to-date answers.
Mathematical reasoning
Grok-4.20 is strong on mathematical reasoning, scoring competitively with GPT-5.5 on competitive mathematics. For pure mathematical work — solving equations, proving theorems, competition problems — Grok-4.20 is among the best models available.
Wit and creativity
Grok-4.20's personality isn't just about humor — it's about creativity. It can write in different styles, generate creative content, and approach problems from unexpected angles. For creative writing, brainstorming, or any task that benefits from lateral thinking, Grok-4.20's unique perspective is an asset.
5. Where Grok-4.20 falls short
Agentic tasks
While Grok-4.20 is in the top 5 on GDPval-AA v2, it's behind Claude Opus 4.8 and GPT-5.5 on complex, multi-step real-world tasks. For the most demanding agentic applications — writing complex software, managing long research projects — Claude is still the better choice.
Factual accuracy
Grok-4.20 has a moderate hallucination rate — higher than Claude (35.9%) and Gemini. Its willingness to answer any question, while valuable for information access, also means it's more likely to confidently state incorrect information. For applications where factual accuracy is critical (research, law, medicine), Claude is more reliable.
Emotional intelligence
Grok-4.20's EQ (emotional intelligence) is estimated at ~110 — lower than Claude (~132), GPT-5.5 (~120), and Gemini (~115). Its irreverent personality, while entertaining, can come across as insensitive in contexts that require empathy. For therapy apps, customer service, or sensitive human interactions, Claude is the better choice.
Safety and reliability
Grok-4.20's uncensored approach, while valuable for information access, also means it's less safe for applications that require careful content moderation. It may generate content that other models would refuse, which can be a liability in professional or regulated environments.
Cost
At 3/15 per 1M tokens, Grok-4.20 is cheaper than Claude (5/25) and GPT-5.5 (5/30) but more expensive than Gemini (2/12) and significantly more expensive than DeepSeek (0.44/0.87). For cost-sensitive applications, Gemini or DeepSeek may be better choices.
6. Grok-4.20 vs other models
| Model | Intelligence Index | Mensa Norway IQ | AI IQ est. | Hallucination | Cost/1M | Personality |
|---|---|---|---|---|---|---|
| Grok-4.20 | top 5 | 145 | ~140 | moderate | 3/15 | irreverent |
| Claude Opus 4.8 | 55.7 | ~130 | ~132 | 35.9% (lowest) | 5/25 | neutral |
| GPT-5.5 | 54.8 | ~145 | ~136 | moderate | 5/30 | neutral |
| Gemini 3.1 Pro | 46.5 | 141 | ~131 | low | 2/12 | neutral |
| DeepSeek V4 Pro | 44.3 | ~111 | ~111 | moderate | 0.44/0.87 | neutral |
The picture: Grok-4.20 is the raw IQ champion — tied for highest IQ, highest composite IQ, and the only model with a real personality. Claude is the capable professional — best at real-world tasks and factual accuracy. GPT-5.5 is the test genius — best at math and visual reasoning. Gemini is the value champion — best cost-effectiveness. And DeepSeek offers remarkable capability at the lowest cost.
For users who want the smartest AI that's also the most interesting to talk to, Grok-4.20 is the clear choice.
7. What does this mean for humans?
Grok-4.20 operates at an IQ of ~145 — smarter than 99.9% of humans. But:
- IQ is narrow: it measures pattern recognition and reasoning, not wisdom, creativity, or judgment
- Intelligence ≠ reliability: Grok-4.20 is the smartest on IQ tests but not the most reliable for factual accuracy
- Personality matters: Grok-4.20 proves that AI can be both highly intelligent and genuinely entertaining
- Knowledge ≠ understanding: Grok-4.20 has access to vast knowledge, but doesn't truly "understand" the way a human does
- No consciousness: high IQ doesn't mean sentience. Grok-4.20 is a sophisticated pattern-matching system, not a thinking being
- The gap is closing: each generation of AI narrows the gap with the smartest humans
The most honest answer: Grok-4.20 is smarter than virtually all humans on cognitive tests, and it's the most distinctive AI model available — one that combines genius-level intelligence with a personality that makes it genuinely engaging to interact with.
Conclusion
- #1 on VexoWire AI IQ composite (~140) — the highest composite IQ among all AI models.
- Unique personality — irreverent, humorous, uncensored. The only frontier model with a real personality.
- Best for: raw IQ, visual reasoning, mathematical reasoning, personality and humor, uncensored responses, real-time information.
- Not best for: agentic tasks (Claude wins), factual accuracy (Claude wins), emotional intelligence (Claude wins), safety and reliability (Claude wins).
- The paradox: Grok-4.20 proves that being funny doesn't mean being less intelligent — it's both the smartest and the most entertaining AI model.
- The lesson: intelligence comes in many forms. Grok-4.20 is the genius with personality — proving that AI can be both brilliant and engaging.


