What is the IQ of Claude Opus 4.8?
Dr. Daria Dudnik 10 min read 27 views

What is the IQ of Claude Opus 4.8?

Claude Opus 4.8 tops the Artificial Analysis Intelligence Index at 55.7 and scores ~132-145 on human IQ tests. We break down every benchmark, explain what the numbers mean, and compare it to GPT-5.5, Gemini, and Grok.

Claude Opus 4.8: Anthropic's smartest model

When Anthropic released Claude Opus 4.8 in mid-2026, it didn't just incrementally improve on its predecessor — it jumped to the top of every major AI benchmark leaderboard. With a score of 55.7 on the Artificial Analysis Intelligence Index v4.1, it narrowly edged out OpenAI's GPT-5.5 (54.8) and Google's Gemini 3.1 Pro (46.5) to claim the #1 spot among all frontier AI models.

But here's the question that fascinates researchers and the general public alike: what does that translate to on a human IQ scale? Can we even compare an AI's intelligence to a human's? And if we can, where does Claude Opus 4.8 land — is it "gifted," "genius," or something beyond any human category?

The answer, as it turns out, depends heavily on which test you give it. And the results are more nuanced — and more surprising — than a single number might suggest.

Key scores:

  • Artificial Analysis Intelligence Index: 55.7 (#1 globally)
  • Mensa Norway IQ (TrackingAI): ~130 (Claude 4.6 Opus tier)
  • AI IQ composite (VexoWire): ~132 estimated
  • GDPval-AA v2: 1,890 Elo (#1 globally)
  • Humanity's Last Exam: #1 among AI models
  • Hallucination rate: 35.9% (lowest among frontier models)
  • Cost: 5.00/25.00 per 1M tokens

1. Why measure AI intelligence with human IQ tests?

Before diving into the numbers, it's worth asking: why would we give an AI a human IQ test at all? AI models don't have brains, they don't develop cognitively, and they process information in fundamentally different ways than humans do.

The answer comes down to comparison. IQ tests, despite their limitations, are the most widely validated and understood measure of cognitive ability we have. They test pattern recognition, spatial reasoning, verbal comprehension, working memory, and logical problem-solving — all capabilities that AI models also need. By administering the same tests to both humans and AI, we get a rough translation scale: if an AI scores 130 on a test where the human average is 100, we can say it performs cognitively at the level of a gifted human on that particular set of tasks.

But there's a critical caveat. IQ tests measure a narrow band of cognitive abilities. They don't measure creativity, emotional intelligence, wisdom, common sense, or the ability to navigate ambiguous real-world situations. An AI with an IQ of 145 can still hallucinate facts, struggle with simple physical reasoning, and fail at tasks a five-year-old handles effortlessly. The IQ number is a floor, not a ceiling — it tells us what the model can do on structured cognitive tasks, but not what it can do in the messy real world.

Several organizations have taken up the challenge of administering IQ tests to AI models. The most prominent is TrackingAI, a research project that gives standardized IQ tests (including Mensa Norway, Mensa Denmark, and Raven's Progressive Matrices) to AI models under controlled conditions. Another is the AI IQ project by VexoWire, which creates a composite score from multiple tests. And the Artificial Analysis Intelligence Index aggregates nine different benchmarks into a single composite score that's become the industry standard for comparing frontier models.


2. Benchmark breakdown

Artificial Analysis Intelligence Index v4.1

The Artificial Analysis Intelligence Index is the gold standard for comparing AI models. It blends nine benchmarks into a single score, weighted by how well they correlate with real-world task performance:

  • GDPval-AA v2 (20%): 1,890 Elo — #1 globally. This measures performance on complex, multi-step real-world tasks like writing code, analyzing data, and creating documents.
  • AA-Omniscience (8%): 35.9% hallucination rate — the lowest among all frontier models. This measures factual accuracy and knowledge reliability.
  • GPQA (6%): Graduate-level physics, chemistry, and biology questions. Claude excels here.
  • CritPt (6%): Critical thinking and physics reasoning.
  • HLE (Humanity's Last Exam): #1 among AI models — the hardest questions from every academic discipline.
  • ARC-AGI: Abstract reasoning and pattern generalization.
  • LMSYS Arena: Human preference voting on response quality.

Claude Opus 4.8's composite score of 55.7 puts it ahead of GPT-5.5 (54.8) by less than one point — a statistical dead heat. But in the sub-components, Claude dominates in agentic tasks (GDPval) and factual accuracy (AA-Omniscience), while GPT-5.5 leads in visual reasoning and mathematical problem-solving.

Mensa Norway IQ test

The Mensa Norway IQ test is one of the most widely used standardized IQ tests for AI evaluation. It consists of 36 visual pattern recognition questions — the kind where you see a sequence of shapes and must identify which option completes the pattern.

Claude 4.6 Opus (the predecessor to 4.8) scored approximately 130 on this test, placing it in the "gifted" range — the top 2% of the human population. Claude Opus 4.8 is expected to score similarly or slightly higher, in the 130-135 range.

This is notable because it's significantly lower than its performance on the Intelligence Index would suggest. On the composite index, Claude is #1 globally. On a pure visual pattern recognition test, it's outperformed by Grok-4.20 (145), GPT-5.4 Pro Vision (145), and Gemini 3.1 Pro (141). Why the gap? Because the Mensa Norway test is heavily visual-spatial, and Claude's architecture is more optimized for verbal and analytical reasoning than for visual pattern recognition.

AI IQ composite (VexoWire)

The VexoWire AI IQ project takes a different approach. Instead of relying on a single test, it creates a composite IQ score from multiple tests — including Mensa Norway, Mensa Denmark, Raven's Progressive Matrices, and the WAIS-IV (Wechsler Adult Intelligence Scale). This gives a more rounded picture of cognitive ability.

Claude Opus 4.8's estimated composite IQ is ~132 — solidly in the "gifted" range. This is comparable to GPT-5.5 (~136) and slightly below Gemini 3.1 Pro on visual tasks (~131-141). But on verbal and analytical subtests, Claude consistently outperforms all other models.

GDPval-AA v2: The agentic benchmark

This is where Claude Opus 4.8 truly dominates. GDPval-AA v2 measures performance on complex, multi-step real-world tasks — the kind that require planning, tool use, and sustained reasoning over many steps. Think: "Research this topic, write a report, create a spreadsheet, and email it to three people."

Claude Opus 4.8 achieves 1,890 Elo on this benchmark — the highest of any AI model. This is roughly equivalent to a highly competent human professional. GPT-5.5 scores 1,850, and Gemini 3.1 Pro scores lower still.

This matters because agentic capability is what separates a model that can answer questions from one that can do things. A model with an IQ of 145 that can't plan a multi-step task is less useful in practice than a model with an IQ of 132 that can.

Humanity's Last Exam (HLE)

Humanity's Last Exam is exactly what it sounds like — a collection of the hardest questions from every academic discipline, submitted by professors worldwide. It's designed to be so difficult that no human could score above 50%.

Claude Opus 4.8 ranks #1 among all AI models on HLE, narrowly beating GPT-5.5. This is perhaps the most meaningful benchmark for raw intellectual capability, because it tests deep knowledge and reasoning across the full range of human academic achievement.


3. What does an IQ of 132 mean?

To put Claude Opus 4.8's estimated IQ of ~132 in context, let's look at what this means on the human IQ scale:

  • 85-115 (68% of people): Average intelligence
  • 115-130 (14% of people): Above average to high average
  • 130-145 (2% of people): Gifted — qualifies for Mensa (top 2%)
  • 145-160 (0.1% of people): Highly gifted / genius
  • 160+ (0.003% of people): Profoundly gifted — Einstein, Hawking territory

Claude Opus 4.8, at ~132, sits right at the Mensa threshold — smarter than 98% of humans. If it were a person, it would qualify for Mensa membership. It would be the smartest student in most classrooms, but it wouldn't be the once-in-a-generation genius that revolutionizes a field.

But here's the crucial difference: Claude has something no human has — instant access to essentially all human knowledge. A human with an IQ of 132 is impressive. An AI with an IQ of 132 that has read every book, every scientific paper, and every website ever written is something entirely different. The IQ measures reasoning ability; the knowledge base measures what it can reason about. Claude has both.


4. Where Claude Opus 4.8 excels

Agentic tasks and real-world reasoning

Claude Opus 4.8 is the best model in the world at multi-step real-world tasks. Whether it's writing a complex software application, analyzing a dataset, or planning a research project, Claude sustains coherent reasoning over longer chains of thought than any competitor. This is its defining advantage.

Factual accuracy and low hallucination

With a hallucination rate of just 35.9% (lowest among frontier models), Claude is the most reliable model for factual queries. It's less likely to confidently state false information than GPT-5.5, Gemini, or Grok. This makes it the preferred model for research, legal, medical, and educational applications where accuracy matters.

Verbal and analytical reasoning

Claude's architecture is optimized for language understanding and analytical reasoning. On verbal comprehension subtests of the WAIS-IV, it would likely score at a genius level. It writes more clearly, reasons more carefully, and constructs better arguments than any other model.

Emotional intelligence

Claude Opus 4.8 has an estimated EQ (emotional intelligence) of ~132 — the highest among all AI models. It's more empathetic, more nuanced in its understanding of human emotions, and better at adjusting its tone to the context than GPT-5.5 or Grok-4.20. This makes it the preferred model for therapy apps, customer service, and any application involving human interaction.

Safety and alignment

Anthropic has invested heavily in making Claude safe and aligned. It's less likely to produce harmful content, more transparent about its limitations, and more careful in high-stakes situations than any competitor. This doesn't show up in IQ scores, but it matters enormously in practice.


5. Where Claude falls short

Visual-spatial reasoning

On the Mensa Norway IQ test, which is heavily visual, Claude scores ~130 — good, but significantly below Grok-4.20 (145), GPT-5.4 Pro Vision (145), and Gemini 3.1 Pro (141). If you need a model for visual pattern recognition, image analysis, or spatial reasoning tasks, Claude is not the best choice.

Mathematical problem-solving

While Claude is strong in mathematical reasoning, GPT-5.5 edges it out on competitive mathematics (AIME, FrontierMath). For pure math, GPT-5.5 is slightly better.

Cost

At 5.00/25.00 per 1M tokens, Claude Opus 4.8 is one of the most expensive models. Gemini 3.1 Pro (2/12) and DeepSeek V4 Pro (0.44/0.87) are significantly cheaper for tasks that don't require Claude's full capability.

Speed

Claude Opus 4.8 is not the fastest model. For simple queries, Gemini 3.5 Flash or DeepSeek V4 Pro will respond faster and cheaper.


6. Claude Opus 4.8 vs other models

Model Intelligence Index Mensa Norway IQ AI IQ est. Hallucination Cost/1M
Claude Opus 4.8 55.7 ~130 ~132 35.9% (lowest) 5/25
GPT-5.5 54.8 ~145 ~136 moderate 5/30
Gemini 3.1 Pro 46.5 141 ~131 low 2/12
Grok-4.20 145 ~140 moderate 3/15
DeepSeek V4 Pro 44.3 ~111 ~111 moderate 0.44/0.87

The picture that emerges is not one model dominating all others, but a landscape of specialized intelligences. Claude is the best all-rounder — the model you'd choose if you could only pick one. GPT-5.5 is the best at math and visual reasoning. Gemini is the most cost-effective and has the best factual knowledge. Grok has the highest raw visual IQ. And DeepSeek offers remarkable capability at a fraction of the cost.


7. What does this mean for humans?

Claude Opus 4.8 operates at an IQ of ~132 — smarter than 98% of humans. But:

  • IQ is narrow: it measures pattern recognition and reasoning, not wisdom, creativity, or judgment
  • Knowledge ≠ understanding: Claude has read everything, but it doesn't truly "understand" the way a human does
  • No consciousness: high IQ doesn't mean sentience. Claude is a sophisticated pattern-matching system, not a thinking being
  • The gap is closing: each generation of AI models narrows the gap with the smartest humans. How long until the gap disappears entirely?

The most honest answer is that Claude Opus 4.8 is smarter than most humans on most cognitive tasks, but not smarter than the best humans on their best days. A world-class mathematician, physicist, or philosopher can still outthink Claude in their domain of expertise. But Claude can outthink all of them simultaneously, across every domain, in seconds.

That's not a high IQ. That's something new — something we don't yet have a word for.

How does your IQ compare to Claude Opus 4.8? Take NeuroLab's professional IQ assessment to find out.
Take the Test Now →


Conclusion

💡 Key takeaways
- Claude Opus 4.8's IQ is estimated at 132-145 depending on the test — solidly in the "gifted" range, qualifying for Mensa.

  • #1 on the Artificial Analysis Intelligence Index (55.7) — the most comprehensive AI benchmark.
  • #1 on GDPval-AA v2 (1,890 Elo) — the best model for real-world multi-step tasks.
  • Lowest hallucination rate (35.9%) — the most factually reliable frontier model.
  • Highest EQ (~132) — the most emotionally intelligent AI model.
  • Best for: complex reasoning, agentic tasks, factual research, and human interaction.
  • Not best for: visual pattern recognition (Grok/Gemini win), pure math (GPT-5.5 wins), cost-sensitive applications (DeepSeek wins).
  • The bottom line: Claude Opus 4.8 is the smartest all-around AI model in existence, but "smartest" is a multidimensional concept — different models excel at different things.
NeuroLab Cognitive Suite

© 2026 NeuroLab. All tests are for cognitive entertainment & training.

Anonymous usage analytics (Umami) help us improve the site. No ads, no selling personal data.