What is the IQ of GPT-5.5?
GPT-5.5 scores 54.8 on the Artificial Analysis Intelligence Index and an estimated IQ of 136 on the AI IQ scale. We break down every benchmark and what it means.
GPT-5.5: OpenAI's frontier model
GPT-5.5 is OpenAI's most advanced model, launched in 2026. It scores 54.8 on the Artificial Analysis Intelligence Index v4.1 — second only to Claude Opus 4.8 (55.7) by less than one point.
On the AI IQ composite scale, GPT-5.5 is estimated at ~136 IQ — the highest of any model tracked by the AI IQ project.
- AI IQ composite (VexoWire): ~136 (#1 globally)
- Mensa Norway IQ (TrackingAI): GPT-5.4 Pro (Vision) = 145
- GDPval-AA v2: 1,531 Elo (#2 globally)
- Cost per task: 0.99 (vs.1.78 for Claude Opus 4.8)
1. Benchmark breakdown
Artificial Analysis Intelligence Index v4.1
GPT-5.5 scores 54.8 — just 0.9 points behind Claude Opus 4.8:
- GDPval-AA v2 (20%): 1,531 Elo — #2
- Terminal-Bench 2.1 (16%): strong
- τ³-Bench Banking (14%): strong
- Humanity's Last Exam (12%): close to #1
- GPQA (6%), SciCode (8%), AA-Omniscience (8%)
AI IQ composite
The AI IQ project places GPT-5.5 at the peak of the bell curve with an estimated IQ of ~136 — the highest of any model. It's followed by GPT-5.4 (~131) and Claude Opus 4.7 (~132).
Mensa Norway IQ test
GPT-5.4 Pro (Vision) scored 145 on TrackingAI's Mensa Norway test — tied with Grok-4.20 Expert Mode for #1 globally. GPT-5.5 is expected to match or exceed this.
iqscore.io test
ChatGPT (free tier) scored 121 on the 36-question test. GPT-5.5 with its full reasoning capabilities would likely score significantly higher.
2. Where GPT-5.5 excels
Abstract reasoning
GPT-5.5 leads on ARC-AGI-1 and ARC-AGI-2, the notoriously difficult pattern-recognition benchmarks that test fluid intelligence.
Mathematical reasoning
Strong performance on FrontierMath (Tiers 1-4), AIME, and ProofBench — the mathematical reasoning dimension of AI IQ.
Cost efficiency
At 0.99 per task, GPT-5.5 is nearly half the cost of Claude Opus 4.8 (1.78) while scoring within one point. It also uses ~30% fewer turns than Opus 4.8 on GDPval tasks.
Speed
GPT-5.5 completes a task in ~3.7 minutes — nearly twice as fast as Claude Opus 4.8 (6.4 minutes).
3. Where it falls short
Hallucination rate
GPT-5.5 has a higher hallucination rate than Claude Opus 4.8. On AA-Omniscience, it scores lower on accuracy and non-hallucination metrics.
Agentic performance
On GDPval-AA v2, GPT-5.5 scores 1,531 Elo vs. Claude Opus 4.8's 1,890 — a ~67% win rate for Claude in head-to-head comparisons.
Emotional intelligence
On the AI IQ EQ scale, GPT-5.5 lags slightly behind Claude Opus 4.8 (~132 EQ). This matters for conversational quality and user-facing applications.
4. GPT-5.5 vs other models
| Model | Intelligence Index | Est. IQ | Cost/task | Speed/task |
|---|---|---|---|---|
| GPT-5.5 | 54.8 | ~136 | $0.99 | 3.7 min |
| Claude Opus 4.8 | 55.7 | ~132-145 | $1.78 | 6.4 min |
| Gemini 3.1 Pro | 46.5 | ~131 | — | — |
| Grok-4.20 | — | 145 | — | — |
| DeepSeek V4 Pro | 44.3 | ~111 | $0.04 | — |
5. What does this mean for humans?
GPT-5.5 operates at an estimated IQ of 136 — higher than 99% of humans. But:
- IQ 136 = top 1%: Mensa-qualifying, but not "genius" (145+)
- AI IQ is domain-specific: excels at math and abstract reasoning, weaker on spatial
- Speed ≠ depth: fast answers aren't always correct answers
- Training data bias: benchmarks may overlap with training data
Conclusion
- Scores 54.8 on the Intelligence Index — #2, just 0.9 behind Claude Opus 4.8.
- Best intelligence-to-cost ratio: 0.99/task vs.1.78 for Claude.
- Fastest frontier model: 3.7 min/task vs. 6.4 min for Claude.
- Leads on abstract and mathematical reasoning (ARC-AGI, FrontierMath).
- Higher hallucination rate than Claude Opus 4.8.
- GPT-5.4 Pro (Vision) scored 145 on Mensa Norway — GPT-5.5 likely matches.