What is the IQ of Gemini 3.1 Pro?
Gemini 3.1 Pro scores 46.5 on the Intelligence Index and 141 on Mensa Norway. We break down Google's flagship model and its cognitive profile.
Gemini 3.1 Pro: Google DeepMind's flagship
Gemini 3.1 Pro is Google DeepMind's most capable model, released in 2026. It scores 46.5 on the Artificial Analysis Intelligence Index v4.1 — ranking #7 globally but excelling in specific dimensions.
On the Mensa Norway IQ test, Gemini 3.1 Pro Preview scored 141 — the third-highest globally, behind only Grok-4.20 (145) and GPT-5.4 Pro Vision (145).
- Mensa Norway IQ (TrackingAI): 141 (#3 globally)
- AI IQ composite (VexoWire): ~131
- AA-Omniscience: 32.9 (#1 globally — highest knowledge accuracy)
- Cost: 2.00/12.00 per 1M tokens
1. Benchmark breakdown
Artificial Analysis Intelligence Index v4.1
Gemini 3.1 Pro scores 46.5 — lower than Claude Opus 4.8 (55.7) and GPT-5.5 (54.8), but strong in specific areas:
- AA-Omniscience (8%): 32.9 — #1 globally (highest knowledge accuracy)
- GPQA (6%): strong
- CritPt (6%): strong (frontier physics)
- GDPval-AA v2 (20%): moderate
Mensa Norway IQ test
Gemini 3.1 Pro Preview scored 141 on TrackingAI's visual pattern test — an exceptional result. With vision enabled, it scored 132 (non-vision verbalized version: 141).
AI IQ composite
The AI IQ project estimates Gemini 3.1 Pro at ~131 IQ — placing it in the top cluster alongside GPT-5.4 and Claude Opus 4.7.
iqscore.io test
Gemini 3.1 Pro scored 129 on the 36-question IQ test — 33/36 correct, with perfect scores on numerical and logical reasoning.
2. Where Gemini 3.1 Pro excels
Knowledge and factual accuracy
#1 on AA-Omniscience (32.9) — highest accuracy, lowest hallucination. Gemini is the most trustworthy model for factual information.
Visual pattern recognition
Mensa Norway IQ of 141 — third globally. Gemini excels at visual-spatial reasoning, especially with vision enabled.
Numerical and logical reasoning
On the iqscore.io test, Gemini scored 9/9 on both numerical and logical sections — perfect.
Speed
Gemini 3.5 Flash (a lighter variant) completes tasks in 1.6 minutes — the fastest of any model with a score above 46.
Cost-effectiveness
At 2.00/12.00 per 1M tokens, Gemini is cheaper than both Claude Opus 4.8 (5/25) and GPT-5.5 (5/30).
3. Where it falls short
Agentic tasks
Gemini 3.1 Pro scores lower on GDPval-AA v2 compared to Claude and GPT. It's less effective at multi-step real-world tasks.
Overall intelligence index
At 46.5, it trails Claude Opus 4.8 (55.7) by 9 points — a significant gap on composite reasoning.
Spatial reasoning (non-vision)
Without vision, Gemini's Mensa Norway score drops from 141 to 132 — showing heavy reliance on visual processing.
4. Gemini 3.1 Pro vs other models
| Model | Intelligence Index | Mensa Norway IQ | AI IQ est. | Cost/1M tokens |
|---|---|---|---|---|
| Gemini 3.1 Pro | 46.5 | 141 | ~131 | 2/12 |
| Claude Opus 4.8 | 55.7 | ~130 | ~132-145 | 5/25 |
| GPT-5.5 | 54.8 | ~145 | ~136 | 5/30 |
| Grok-4.20 | — | 145 | — | 3/15 |
| DeepSeek V4 Pro | 44.3 | ~111 | ~111 | 0.44/0.87 |
5. What does this mean for humans?
Gemini 3.1 Pro operates at an estimated IQ of 131-141 — higher than 98-99% of humans. But:
- IQ 141 = top 0.25%: higher than most Mensa members
- Knowledge is not reasoning: knowing facts ≠ solving problems
- Visual genius, agentic average: excels at patterns, weaker on multi-step tasks
- Speed advantage: faster and cheaper than Claude or GPT
Conclusion
- Mensa Norway IQ: 141 — #3 globally, exceptional visual reasoning.
- #1 on AA-Omniscience: most factual, lowest hallucination.
- Intelligence Index: 46.5 — lower than Claude/GPT on composite tasks.
- Cheapest frontier model: 2/12 per 1M tokens.
- Best for factual queries and visual patterns, not agentic workflows.
- Perfect 9/9 on numerical and logical reasoning (iqscore.io).