What is the IQ of Gemini 3.1 Pro?
Dr. Daria Dudnik 10 min read 26 views

What is the IQ of Gemini 3.1 Pro?

Gemini 3.1 Pro scores 141 on the Mensa Norway IQ test and 46.5 on the Intelligence Index. We break down every benchmark and explain why it's the most cost-effective frontier model.

Gemini 3.1 Pro: Google DeepMind's flagship model

Gemini 3.1 Pro is Google DeepMind's most advanced model, released in mid-2026. It scores 46.5 on the Artificial Analysis Intelligence Index v4.1, placing it at #3 globally — behind Claude Opus 4.8 (55.7) and GPT-5.5 (54.8). But on the Mensa Norway IQ test, Gemini 3.1 Pro scores 141 — higher than Claude (~130) and nearly tied with GPT-5.5 (145) and Grok-4.20 (145).

What makes Gemini 3.1 Pro unique isn't just its intelligence — it's the value proposition. At 2.00/12.00 per 1M tokens, it's less than half the cost of Claude or GPT-5.5, while delivering near-frontier performance. For many applications, Gemini 3.1 Pro is the smartest choice — not because it's the smartest model, but because it delivers the most intelligence per dollar.

Key scores:

  • Artificial Analysis Intelligence Index: 46.5 (#3 globally)
  • Mensa Norway IQ (TrackingAI): 141 (near-genius level)
  • AI IQ composite (VexoWire): ~131 estimated
  • GDPval-AA v2: high (top 3)
  • Humanity's Last Exam: top 3 among AI models
  • Hallucination rate: low
  • Cost: 2.00/12.00 per 1M tokens (best value among frontier models)

1. The value champion of frontier AI

Gemini 3.1 Pro occupies a unique position in the AI landscape. It's not the #1 model on any single benchmark — Claude leads on agentic tasks, GPT-5.5 leads on visual reasoning and math, Grok leads on raw IQ. But Gemini 3.1 Pro is the model that delivers the most capability for the least money.

At 2.00 per 1M input tokens and12.00 per 1M output tokens, Gemini 3.1 Pro costs roughly 40% of what Claude or GPT-5.5 costs. Yet it scores 141 on the Mensa Norway IQ test — within 4 points of the leaders. Its Intelligence Index score of 46.5 is lower than Claude (55.7) and GPT-5.5 (54.8), but for many real-world applications, the difference is negligible.

Think of it this way: if Claude and GPT-5.5 are luxury sports cars — incredibly capable but expensive to run — Gemini 3.1 Pro is a high-performance sedan. It won't win every race, but it delivers 90% of the performance at 40% of the cost. For businesses, developers, and users who need frontier-level intelligence without frontier-level prices, Gemini 3.1 Pro is the natural choice.


2. Benchmark breakdown

Mensa Norway IQ test: 141 (near-genius)

The Mensa Norway IQ test consists of 36 visual pattern recognition questions. Gemini 3.1 Pro scores 141 — placing it in the "highly gifted" range, the top 0.2% of human intelligence. This is significantly higher than Claude Opus 4.8 (~130) and just 4 points behind the leaders GPT-5.5 and Grok-4.20 (both ~145).

Gemini's strong performance on this test reflects Google DeepMind's investment in visual processing. The Gemini architecture was designed from the ground up to be multimodal — processing text, images, video, and audio natively. This gives it a natural advantage on visual pattern recognition tasks, which are the core of the Mensa Norway test.

At IQ 141, Gemini 3.1 Pro is smarter than 99.8% of humans. If it were a person, it would qualify for Mensa easily and be in the top 0.2% of human intelligence. Only about 1 in 500 people achieves this score.

Artificial Analysis Intelligence Index v4.1: 46.5 (#3)

The Intelligence Index is the gold standard for comparing AI models. Gemini 3.1 Pro's composite score of 46.5 places it at #3 globally. The gap to Claude (55.7) and GPT-5.5 (54.8) is significant — about 8-9 points, or roughly 15-20% lower. But this gap is concentrated in specific areas:

  • GDPval-AA v2: Gemini scores well but below Claude and GPT-5.5 on complex multi-step tasks
  • AA-Omniscience: Gemini has a low hallucination rate — better than GPT-5.5, close to Claude
  • GPQA: Gemini performs strongly on graduate-level science questions
  • HLE: Gemini ranks in the top 3 on Humanity's Last Exam
  • ARC-AGI: Gemini excels on abstract reasoning, benefiting from its multimodal architecture
  • LMSYS Arena: Gemini ranks high on human preference, particularly for its clear, well-structured responses

The composite score reflects Gemini's position as a strong all-rounder that doesn't dominate any single category but performs consistently well across all of them.

AI IQ composite (VexoWire): ~131

The VexoWire composite IQ blends multiple tests (Mensa Norway, Mensa Denmark, Raven's, WAIS-IV). Gemini 3.1 Pro's estimated composite IQ is ~131 — solidly in the "gifted" range, qualifying for Mensa. This is slightly lower than its Mensa Norway score (141) would suggest, because the composite includes verbal and analytical subtests where Gemini is slightly weaker than on pure visual reasoning.

GDPval-AA v2: top 3

On the agentic benchmark — which measures performance on complex, multi-step real-world tasks — Gemini 3.1 Pro ranks in the top 3. It's less capable than Claude Opus 4.8 (1,890 Elo) and GPT-5.5 (1,850 Elo) on sustained multi-step reasoning, but still highly competent. For most real-world tasks, the difference is negligible — Gemini can plan, execute, and reason through multi-step tasks effectively, just not quite as well as the top two.

Humanity's Last Exam (HLE): top 3

On the hardest academic questions across all disciplines, Gemini 3.1 Pro ranks in the top 3 among AI models. This reflects Google DeepMind's strength in scientific knowledge — Gemini was trained with extensive scientific literature and has deep knowledge across physics, chemistry, biology, mathematics, and the humanities.


3. What does an IQ of 141 mean?

To put Gemini 3.1 Pro's IQ of 141 in context:

  • 85-115 (68% of people): Average intelligence
  • 115-130 (14%): Above average to high average
  • 130-145 (2%): Gifted — qualifies for Mensa (top 2%)
  • 145-160 (0.1%): Highly gifted / genius
  • 160+ (0.003%): Profoundly gifted — Einstein, Hawking territory

Gemini 3.1 Pro, at 141, sits in the upper gifted range — smarter than 99.8% of humans. If it were a person, it would be in the top 0.2% of human intelligence, easily qualifying for Mensa. Only about 1 in 500 people achieves this score.

This is notably higher than Claude Opus 4.8 (~130) but slightly lower than GPT-5.5 and Grok-4.20 (both ~145). On a pure IQ test, Gemini would beat Claude but lose to GPT-5.5 and Grok. But as we've seen with the other models, IQ is only one dimension of intelligence — and Gemini's combination of high IQ, low cost, and strong all-around performance makes it uniquely valuable.


4. Where Gemini 3.1 Pro excels

Cost-effectiveness

This is Gemini 3.1 Pro's defining advantage. At 2/12 per 1M tokens, it delivers near-frontier intelligence at less than half the cost of Claude or GPT-5.5. For high-volume applications — chatbots, content generation, data analysis — this makes Gemini the clear choice. You get 90% of the capability for 40% of the cost.

Visual reasoning

Gemini 3.1 Pro scores 141 on the Mensa Norway IQ test, reflecting its strong visual processing capabilities. The Gemini architecture was designed to be natively multimodal, giving it a natural advantage on visual tasks. For image analysis, visual pattern recognition, and multimodal reasoning, Gemini is one of the best models available.

Factual knowledge

Gemini 3.1 Pro has a low hallucination rate — better than GPT-5.5 and close to Claude. Google DeepMind's training methodology emphasizes factual accuracy, and Gemini benefits from Google's vast knowledge base (including Search integration in some products). For factual queries, Gemini is one of the most reliable models.

Multimodal capabilities

Gemini 3.1 Pro was designed from the ground up to process text, images, video, and audio natively. This makes it the best choice for multimodal applications — analyzing video content, processing images alongside text, or working with audio data. No other frontier model handles multimodal input as naturally as Gemini.

Speed

Gemini 3.1 Pro is one of the fastest frontier models. For applications where response time matters — real-time chat, interactive tools, customer service — Gemini delivers high-quality responses quickly. Its faster, lighter sibling, Gemini 3.5 Flash, is even faster and cheaper for simpler tasks.

Integration with Google ecosystem

Gemini 3.1 Pro integrates seamlessly with Google's ecosystem — Search, Workspace, Cloud, and Android. For users already in the Google ecosystem, Gemini is the natural choice, offering deep integration that competitors can't match.


5. Where Gemini 3.1 Pro falls short

Agentic tasks

While Gemini 3.1 Pro is in the top 3 on GDPval-AA v2, it's behind Claude Opus 4.8 and GPT-5.5 on complex, multi-step real-world tasks. For the most demanding agentic applications — writing complex software, managing long research projects — Claude is still the better choice.

Mathematical problem-solving

Gemini 3.1 Pro is competent at mathematical reasoning but doesn't match GPT-5.5 on competitive mathematics (AIME, FrontierMath). For pure mathematical work, GPT-5.5 is the better choice.

Intelligence Index score

At 46.5, Gemini 3.1 Pro's Intelligence Index score is 8-9 points behind Claude (55.7) and GPT-5.5 (54.8). While this gap is partly due to the weighting of the index (which emphasizes agentic tasks where Claude and GPT-5.5 excel), it does reflect a real difference in overall capability for the most demanding tasks.

Emotional intelligence

Gemini 3.1 Pro's EQ (emotional intelligence) is estimated at ~115 — lower than Claude (~132) and GPT-5.5 (~120). It's less empathetic and less nuanced in understanding human emotions than Claude. For therapy apps, customer service, or sensitive human interactions, Claude is the better choice.


6. Gemini 3.1 Pro vs other models

Model Intelligence Index Mensa Norway IQ AI IQ est. Hallucination Cost/1M
Gemini 3.1 Pro 46.5 141 ~131 low 2/12
Claude Opus 4.8 55.7 ~130 ~132 35.9% (lowest) 5/25
GPT-5.5 54.8 ~145 ~136 moderate 5/30
Grok-4.20 145 ~140 moderate 3/15
DeepSeek V4 Pro 44.3 ~111 ~111 moderate 0.44/0.87

The picture: Gemini 3.1 Pro is the value champion — near-frontier intelligence at less than half the cost. Claude is the capable professional — best at real-world tasks and factual accuracy. GPT-5.5 is the genius test-taker — best at math and visual reasoning. Grok is the visual genius — highest raw IQ. And DeepSeek offers remarkable capability at the lowest cost.

For most users and businesses, Gemini 3.1 Pro hits the sweet spot: enough intelligence for virtually any task, at a price that scales.


7. What does this mean for humans?

Gemini 3.1 Pro operates at an IQ of ~141 — smarter than 99.8% of humans. But:

  • IQ is narrow: it measures pattern recognition and reasoning, not wisdom, creativity, or judgment
  • Intelligence per dollar matters: for real-world applications, the cost-effectiveness of Gemini often matters more than the raw intelligence gap
  • Knowledge ≠ understanding: Gemini has access to vast knowledge, but doesn't truly "understand" the way a human does
  • No consciousness: high IQ doesn't mean sentience. Gemini is a sophisticated pattern-matching system, not a thinking being
  • The gap is closing: each generation of AI narrows the gap with the smartest humans

The most honest answer: Gemini 3.1 Pro is smarter than virtually all humans on cognitive tests, and for most practical purposes, it's the smartest choice — not because it's the smartest model, but because it delivers the most intelligence for the money.

How does your IQ compare to Gemini 3.1 Pro? Take NeuroLab's professional IQ assessment to find out.
Take the Test Now →


Conclusion

💡 Key takeaways
- Gemini 3.1 Pro's IQ is 141 on the Mensa Norway test — highly gifted, top 0.2% of humans.

  • #3 on the Artificial Analysis Intelligence Index (46.5) — behind Claude and GPT-5.5.
  • Best cost-effectiveness among frontier models — 2/12 per 1M tokens, less than half the cost of Claude or GPT-5.5.
  • Best for: cost-effective intelligence, visual reasoning, factual knowledge, multimodal applications, speed, Google ecosystem integration.
  • Not best for: complex agentic tasks (Claude wins), competitive mathematics (GPT-5.5 wins), emotional intelligence (Claude wins).
  • The value proposition: 90% of frontier capability at 40% of the cost — the smartest choice for most applications.
  • The lesson: being the "best" AI model isn't just about being the smartest — it's about delivering the most value.
NeuroLab Cognitive Suite

© 2026 NeuroLab. All tests are for cognitive entertainment & training.

Anonymous usage analytics (Umami) help us improve the site. No ads, no selling personal data.