What is the IQ of Claude Opus 4.8?
Claude Opus 4.8 tops the Artificial Analysis Intelligence Index at 55.7. But what does that translate to on a human IQ scale? We break down the benchmarks, scores, and what they mean.
Claude Opus 4.8: Anthropic's smartest model
Claude Opus 4.8 is Anthropic's flagship AI model, released in mid-2026. It currently leads the Artificial Analysis Intelligence Index v4.1 with a score of 55.7, narrowly ahead of OpenAI's GPT-5.5 at 54.8.
But how does this translate to a human IQ score? The answer depends on which benchmark you use.
- Mensa Norway IQ (TrackingAI): ~130 (Claude 4.6 Opus tier)
- AI IQ composite (VexoWire): ~132 estimated
- GDPval-AA v2: 1,890 Elo (#1 globally)
- Humanity's Last Exam: #1 among AI models
1. Benchmark breakdown
Artificial Analysis Intelligence Index v4.1
Claude Opus 4.8 scores 55.7 on this composite index, which blends nine benchmarks:
- GDPval-AA v2 (20%): 1,890 Elo — #1
- Terminal-Bench 2.1 (16%): strong
- τ³-Bench Banking (14%): strong
- Humanity's Last Exam (12%): #1
- GPQA (6%), SciCode (8%), AA-Omniscience (8%)
Mensa Norway IQ test
On TrackingAI's Mensa Norway benchmark (April 2026), Claude 4.6 Opus scored 130. Opus 4.8 is expected to score 132-135 based on improvement trends.
AI IQ composite
The AI IQ project estimates Claude Opus 4.7 at ~132. Opus 4.8 likely sits at ~134-136.
iqscore.io test
Claude Sonnet 5 scored 145 (perfect 36/36) on the iqscore.io 36-question test. Opus 4.8 would likely match or exceed this.
2. Where Claude Opus 4.8 excels
Agentic tasks
Opus 4.8 leads GDPval-AA v2 at 1,890 Elo — a ~67% win rate against GPT-5.5. This measures real-world knowledge work: research, analysis, multi-step reasoning.
Scientific reasoning
Opus 4.8 leads Humanity's Last Exam, overtaking both OpenAI and Google. It also beats Gemini 3.1 Pro on CritPt (frontier physics).
Low hallucination
Opus 4.8 has a hallucination rate of 35.9% — substantially lower than Google and OpenAI models. It scores 27.4 on AA-Omniscience, #2 behind Gemini 3.1 Pro.
Efficiency
Opus 4.8 uses 15% fewer turns and 35% fewer output tokens than Opus 4.7, while scoring 4 points higher.
3. Where it falls short
Cost
At 5.00/25.00 per 1M tokens (input/output), Opus 4.8 is expensive. DeepSeek V4 Pro delivers similar agentic performance at 0.04 per task vs.1.78 for Opus 4.8.
Speed
Opus 4.8 completes a task in ~6.4 minutes. Gemini 3.5 Flash does it in 1.6 minutes with a score of 50.2.
Spatial reasoning
On the Mensa Norway visual pattern test, Claude models historically score lower than vision-enabled models like GPT-5.4 Pro (145) and Grok-4.20 (145).
4. Claude Opus 4.8 vs other models
| Model | Intelligence Index | Est. IQ | Cost/1M tokens |
|---|---|---|---|
| Claude Opus 4.8 | 55.7 | ~132-145 | 5/25 |
| GPT-5.5 | 54.8 | ~136 | 5/30 |
| Gemini 3.1 Pro | 46.5 | ~131 | 2/12 |
| Grok-4.20 | — | 145 | 3/15 |
| DeepSeek V4 Pro | 44.3 | ~111 | 0.44/0.87 |
5. What does this mean for humans?
Claude Opus 4.8 operates at a cognitive level that exceeds 98% of humans on standard IQ tests. But:
- AI IQ is "jagged": Claude excels at reasoning and analysis but can't tie shoelaces
- Pattern recognition ≠ understanding: it solves puzzles without consciousness
- Benchmark scores can be gamed: training data may include similar problems
- Human IQ is broader: emotional, social, physical intelligence matter
Conclusion
- It leads the Artificial Analysis Intelligence Index at 55.7, ahead of GPT-5.5.
- It's #1 on agentic tasks (GDPval-AA) and Humanity's Last Exam.
- Lowest hallucination rate among frontier models.
- Expensive but capable: best for complex knowledge work.
- AI IQ ≠ human IQ: scores measure specific cognitive dimensions, not general intelligence.