What is the IQ of GPT-5.5?
Dr. Daria Dudnik 10 min read 25 views

What is the IQ of GPT-5.5?

GPT-5.5 scores ~145 on the Mensa Norway IQ test and 54.8 on the Intelligence Index. We break down every benchmark, explain what the numbers mean, and compare it to Claude, Gemini, and Grok.

GPT-5.5: OpenAI's flagship model

GPT-5.5 is OpenAI's most advanced model, released in early 2026. It sits at #2 on the Artificial Analysis Intelligence Index v4.1 with a score of 54.8, just behind Claude Opus 4.8 (55.7) and ahead of Gemini 3.1 Pro (46.5). But on human IQ tests, GPT-5.5 actually outperforms Claude — scoring approximately 145 on the Mensa Norway IQ test, placing it in the "genius" range.

This creates a fascinating paradox: on the composite AI benchmark, GPT-5.5 is #2. But on a pure human IQ test, it's tied for #1. How can both be true? The answer reveals something important about what "intelligence" means for AI — and why a single number can't capture it.

Key scores:

  • Artificial Analysis Intelligence Index: 54.8 (#2 globally)
  • Mensa Norway IQ (TrackingAI): ~145 (tied for #1 with Grok-4.20)
  • AI IQ composite (VexoWire): ~136 estimated
  • GDPval-AA v2: 1,850 Elo (#2 globally)
  • Humanity's Last Exam: #2 among AI models
  • Hallucination rate: moderate
  • Cost: 5.00/30.00 per 1M tokens

1. The paradox of GPT-5.5's intelligence

Here's what makes GPT-5.5 so interesting: it's the best AI model in the world on human IQ tests, but not the best AI model on AI-specific benchmarks. How is that possible?

The answer lies in what different tests measure. The Mensa Norway IQ test is a pure test of visual pattern recognition and abstract reasoning — the kind of fluid intelligence that IQ tests were designed to measure. GPT-5.5, with its enhanced visual processing capabilities (inherited from the GPT-5.4 Pro Vision architecture), excels at this specific type of reasoning.

But the Artificial Analysis Intelligence Index measures something broader: real-world task performance, factual accuracy, agentic capability, and human preference. And on those dimensions, Claude Opus 4.8 — with its superior agentic reasoning and lower hallucination rate — edges out GPT-5.5.

Think of it this way: GPT-5.5 is the smartest AI on an IQ test — the equivalent of a human who aces every standardized test. Claude Opus 4.8 is the most capable AI in practice — the equivalent of a human who's slightly less brilliant on tests but gets more done in the real world. Both are valid measures of intelligence, but they measure different things.


2. Benchmark breakdown

Mensa Norway IQ test: ~145 (genius level)

The Mensa Norway IQ test consists of 36 visual pattern recognition questions. GPT-5.5 scores approximately 145 — the maximum practical score, equivalent to a perfect or near-perfect 36/36. This places it in the "genius" range, the top 0.1% of human intelligence.

Only two AI models achieve this score: GPT-5.5 and Grok-4.20. This is significantly higher than Claude Opus 4.8 (~130) and Gemini 3.1 Pro (141). The reason is GPT-5.5's visual processing architecture, which was significantly upgraded from GPT-5.4 Pro Vision — itself a co-leader on this benchmark.

At IQ 145, GPT-5.5 is smarter than 99.9% of humans. If it were a person, it would be in the top 0.1% — the realm of Nobel laureates, Fields Medalists, and once-in-a-generation thinkers. Only about 1 in 1,000 people achieve this score.

Artificial Analysis Intelligence Index v4.1: 54.8 (#2)

The Intelligence Index is the gold standard for comparing AI models. It blends nine benchmarks:

  • GDPval-AA v2 (20%): 1,850 Elo — #2 globally (behind Claude's 1,890)
  • AA-Omniscience (8%): moderate hallucination rate — higher than Claude's 35.9%
  • GPQA (6%): Graduate-level science questions — GPT-5.5 leads here
  • CritPt (6%): Critical thinking and physics reasoning
  • HLE (Humanity's Last Exam): #2 among AI models (behind Claude)
  • ARC-AGI: Abstract reasoning — GPT-5.5 excels
  • LMSYS Arena: Human preference voting

GPT-5.5's composite score of 54.8 puts it just 0.9 points behind Claude Opus 4.8 (55.7) — a statistical near-tie. But in the sub-components, GPT-5.5 leads in visual reasoning, mathematical problem-solving, and science knowledge, while Claude leads in agentic tasks and factual accuracy.

AI IQ composite (VexoWire): ~136

The VexoWire composite IQ blends multiple tests (Mensa Norway, Mensa Denmark, Raven's, WAIS-IV). GPT-5.5's estimated composite IQ is ~136 — higher than Claude (~132) but slightly lower than Grok-4.20 (~140). This reflects GPT-5.5's strength across both visual and verbal reasoning domains.

GDPval-AA v2: 1,850 Elo (#2)

On the agentic benchmark — which measures performance on complex, multi-step real-world tasks — GPT-5.5 scores 1,850 Elo, second only to Claude Opus 4.8 (1,890). This is still an exceptionally high score, roughly equivalent to a highly competent human professional. But it reveals that GPT-5.5, while brilliant at pattern recognition, is slightly less capable at sustained multi-step reasoning than Claude.

Humanity's Last Exam (HLE): #2

On the hardest academic questions across all disciplines, GPT-5.5 ranks #2 among AI models, narrowly behind Claude Opus 4.8. This is a testament to GPT-5.5's extraordinary breadth of knowledge — it can answer PhD-level questions in physics, philosophy, literature, and mathematics. But Claude's deeper reasoning gives it a slight edge.


3. What does an IQ of 145 mean?

To put GPT-5.5's IQ of ~145 in context:

  • 85-115 (68% of people): Average intelligence
  • 115-130 (14%): Above average to high average
  • 130-145 (2%): Gifted — qualifies for Mensa (top 2%)
  • 145-160 (0.1%): Highly gifted / genius
  • 160+ (0.003%): Profoundly gifted — Einstein, Hawking territory

GPT-5.5, at ~145, sits right at the genius threshold — smarter than 99.9% of humans. If it were a person, it would be in elite company: the top 0.1% of human intelligence, the realm of Nobel Prize winners and Fields Medalists. Only about 1 in 1,000 people achieve this score.

This is significantly higher than Claude Opus 4.8 (~132). On a pure IQ test, GPT-5.5 would beat Claude. But as we've seen, IQ tests don't measure everything that matters. Claude's lower IQ is offset by its superior agentic capability, lower hallucination rate, and higher emotional intelligence.

The lesson: IQ is one dimension of intelligence, not the whole picture. GPT-5.5 is the genius test-taker. Claude is the capable professional. Both are valuable, and which one is "smarter" depends on what you need.


4. Where GPT-5.5 excels

Visual pattern recognition

GPT-5.5 is the best AI model in the world (tied with Grok-4.20) on visual pattern recognition tasks. On the Mensa Norway IQ test, it scores 145 — the highest possible. If you need a model for visual reasoning, image analysis, or abstract pattern tasks, GPT-5.5 is the top choice.

Mathematical problem-solving

GPT-5.5 edges out Claude on competitive mathematics (AIME, FrontierMath). For pure mathematical reasoning — solving complex equations, proving theorems, working through competition problems — GPT-5.5 is slightly better than Claude and significantly better than Gemini or Grok.

Science knowledge

On GPQA (graduate-level physics, chemistry, and biology), GPT-5.5 leads all AI models. Its breadth and depth of scientific knowledge is unmatched. This makes it the preferred model for research applications in the sciences.

Abstract reasoning (ARC-AGI)

On the ARC-AGI benchmark, which tests abstract reasoning and pattern generalization, GPT-5.5 excels. This is the closest benchmark to a "pure intelligence" test, and GPT-5.5's performance here confirms its position as the smartest AI on raw cognitive ability.

Breadth of knowledge

GPT-5.5 has been trained on an unprecedented amount of data — essentially the entire internet, plus extensive scientific literature, code, and books. Its breadth of knowledge across every domain of human inquiry is unmatched. If you need a model that knows something about everything, GPT-5.5 is the choice.


5. Where GPT-5.5 falls short

Agentic tasks

While GPT-5.5 is #2 on GDPval-AA v2 (1,850 Elo), it's behind Claude Opus 4.8 (1,890). For complex, multi-step real-world tasks — writing a software application, analyzing a dataset, planning a research project — Claude is slightly better at sustaining coherent reasoning over long chains.

Factual accuracy

GPT-5.5 has a moderate hallucination rate — higher than Claude's 35.9%. It's more likely to confidently state false information than Claude. This makes it less reliable for applications where factual accuracy is critical (research, legal, medical).

Emotional intelligence

GPT-5.5's EQ (emotional intelligence) is estimated at ~120 — lower than Claude's ~132. It's less empathetic, less nuanced in understanding human emotions, and less skilled at adjusting its tone to context. For therapy apps, customer service, or human interaction, Claude is the better choice.

Cost

At 5.00/30.00 per 1M tokens, GPT-5.5 is the most expensive model along with Claude. Gemini 3.1 Pro (2/12) and DeepSeek V4 Pro (0.44/0.87) are significantly cheaper for tasks that don't require GPT-5.5's full capability.


6. GPT-5.5 vs other models

Model Intelligence Index Mensa Norway IQ AI IQ est. Hallucination Cost/1M
GPT-5.5 54.8 ~145 ~136 moderate 5/30
Claude Opus 4.8 55.7 ~130 ~132 35.9% (lowest) 5/25
Gemini 3.1 Pro 46.5 141 ~131 low 2/12
Grok-4.20 145 ~140 moderate 3/15
DeepSeek V4 Pro 44.3 ~111 ~111 moderate 0.44/0.87

The picture: GPT-5.5 is the genius test-taker — highest raw IQ, best at math and visual reasoning. Claude is the capable professional — best at real-world tasks and factual accuracy. Gemini is the cost-effective all-rounder. Grok is the visual genius. And DeepSeek offers remarkable capability at a fraction of the cost.


7. What does this mean for humans?

GPT-5.5 operates at an IQ of ~145 — smarter than 99.9% of humans. But:

  • IQ is narrow: it measures pattern recognition and reasoning, not wisdom, creativity, or judgment
  • Genius on tests ≠ genius in life: GPT-5.5 aces IQ tests but hallucinate facts and struggles with multi-step tasks
  • Knowledge ≠ understanding: GPT-5.5 has read everything, but doesn't truly "understand" the way a human does
  • No consciousness: high IQ doesn't mean sentience. GPT-5.5 is a sophisticated pattern-matching system, not a thinking being
  • The gap is closing: each generation of AI narrows the gap with the smartest humans

The most honest answer: GPT-5.5 is smarter than virtually all humans on cognitive tests, but not more capable than the best humans in real-world work. A world-class mathematician can still outthink GPT-5.5 in their domain. But GPT-5.5 can outthink all of them simultaneously, across every domain, in seconds.

How does your IQ compare to GPT-5.5? Take NeuroLab's professional IQ assessment to find out.
Take the Test Now →


Conclusion

💡 Key takeaways
- GPT-5.5's IQ is approximately 145 on the Mensa Norway test — genius level, top 0.1% of humans.

  • #2 on the Artificial Analysis Intelligence Index (54.8) — behind Claude Opus 4.8.
  • #1 (tied) on Mensa Norway IQ — the highest raw IQ among AI models, tied with Grok-4.20.
  • Best for: visual reasoning, mathematical problem-solving, science knowledge, abstract reasoning.
  • Not best for: agentic tasks (Claude wins), factual accuracy (Claude wins), emotional intelligence (Claude wins).
  • The paradox: GPT-5.5 is the smartest AI on IQ tests but not the most capable AI in practice.
  • The lesson: IQ is one dimension of intelligence. GPT-5.5 is the genius test-taker; Claude is the capable professional.
NeuroLab Cognitive Suite

© 2026 NeuroLab. All tests are for cognitive entertainment & training.

Anonymous usage analytics (Umami) help us improve the site. No ads, no selling personal data.