What is the IQ of DeepSeek V4?
Dr. Daria Dudnik 10 min read 0 views

What is the IQ of DeepSeek V4?

DeepSeek V4 Pro scores 111 on the Mensa Norway IQ test and 44.3 on the Intelligence Index. We break down every benchmark and explain why it's the best budget AI model.

DeepSeek V4 Pro: China's open-source champion

DeepSeek V4 Pro is the flagship model from DeepSeek, the Chinese AI lab that shocked the world with its combination of low cost and high performance. Released in early 2026, DeepSeek V4 Pro scores 111 on the Mensa Norway IQ test and 44.3 on the Artificial Analysis Intelligence Index v4.1 — placing it in the top 10 globally, but well behind the frontier models (Claude at 55.7, GPT-5.5 at 54.8).

What makes DeepSeek V4 Pro remarkable is not its raw intelligence — it's its price. At 0.44/0.87 per 1M tokens, it costs roughly 10x less than Claude (5/25) and GPT-5.5 (5/30), and 5x less than Gemini (2/12). For the price of a single query to Claude, you could run ten queries to DeepSeek V4 Pro. And yet, it scores 111 on an IQ test — higher than approximately 25% of humans.

DeepSeek V4 Pro is the people's AI: a model that doesn't compete at the frontier of intelligence, but brings capable AI to everyone at a price that's practically free. For students, hobbyists, small businesses, and developers in cost-sensitive markets, DeepSeek V4 Pro is often the only viable option.

Key scores:

  • Artificial Analysis Intelligence Index: 44.3 (top 10 globally)
  • Mensa Norway IQ (TrackingAI): 111 (above average human)
  • AI IQ composite (VexoWire): ~111 estimated
  • GDPval-AA v2: top 10
  • Humanity's Last Exam: top 10
  • Hallucination rate: moderate
  • Cost: 0.44/0.87 per 1M tokens (lowest among all major models)
  • Open-source: weights available for local deployment

1. The budget revolution

DeepSeek V4 Pro represents a fundamental shift in the AI landscape: the democratization of intelligence. While American labs (OpenAI, Anthropic, Google, xAI) compete at the frontier — pushing IQ scores from 130 to 140 to 145 — DeepSeek has focused on a different goal: making capable AI affordable to everyone.

At 0.44 per 1M input tokens and0.87 per 1M output tokens, DeepSeek V4 Pro is staggeringly cheap. To put this in perspective: processing a typical book-length document (about 100,000 tokens) costs approximately 0.04 with input and0.09 with output. The same task would cost 0.50-1.00 with Gemini, 1.00-3.00 with Claude or GPT-5.5.

This isn't just a discount — it's a different category of product. DeepSeek V4 Pro makes AI accessible to:

  • Students who can't afford $20/month subscriptions
  • Small businesses in developing markets
  • Hobbyists and tinkerers running experiments
  • Developers building high-volume applications
  • Researchers processing massive datasets on limited budgets

And the remarkable thing is: DeepSeek V4 Pro is genuinely capable. It scores 111 on the Mensa Norway IQ test — above the human average of 100, and higher than about 25% of people. It's not a genius, but it's smart enough for most everyday tasks: writing, coding, analysis, translation, summarization.


2. Benchmark breakdown

Mensa Norway IQ test: 111 (above average)

The Mensa Norway IQ test consists of 36 visual pattern recognition questions. DeepSeek V4 Pro scores 111 — above the human average of 100, placing it in the "above average" range. This is higher than approximately 25% of humans, but well below the frontier models (Claude ~130, Gemini 141, GPT-5.5 ~145, Grok-4.20 145).

An IQ of 111 means DeepSeek V4 Pro can recognize visual patterns and reason abstractly at a level comparable to a bright college student. It won't solve the hardest puzzles, but it handles most everyday reasoning tasks competently.

The gap between DeepSeek V4 Pro (111) and the frontier models (130-145) is significant — about 20-35 IQ points, or roughly 1-2 standard deviations. This means the frontier models are substantially smarter on raw cognitive tests. But for tasks that don't require genius-level reasoning — which is most tasks — DeepSeek V4 Pro is more than adequate.

Artificial Analysis Intelligence Index v4.1: 44.3 (top 10)

On the Intelligence Index, DeepSeek V4 Pro scores 44.3 — placing it in the top 10 globally, but about 10-11 points behind the frontier leaders (Claude 55.7, GPT-5.5 54.8). This gap is significant but expected given the price difference.

Breaking down the Index components:

  • GDPval-AA v2: DeepSeek scores reasonably but trails Claude and GPT-5.5 on complex multi-step tasks
  • AA-Omniscience: DeepSeek has a moderate hallucination rate — not as low as Claude, but acceptable for most uses
  • GPQA: DeepSeek performs adequately on graduate-level science questions
  • HLE: DeepSeek ranks in the top 10 on Humanity's Last Exam
  • ARC-AGI: DeepSeek struggles with abstract reasoning compared to frontier models
  • LMSYS Arena: DeepSeek gets decent human preference scores, particularly for coding tasks

AI IQ composite (VexoWire): ~111

The VexoWire composite IQ blends multiple tests (Mensa Norway, Mensa Denmark, Raven's, WAIS-IV). DeepSeek V4 Pro's estimated composite IQ is ~111 — consistent with its Mensa Norway score. This places it firmly in the "above average" range, comparable to a bright human but well below genius level.

GDPval-AA v2: top 10

On the agentic benchmark, DeepSeek V4 Pro ranks in the top 10. It's significantly less capable than Claude (1,890 Elo) and GPT-5.5 (1,850 Elo) on complex multi-step tasks, but can handle moderate agentic workloads. For simple automation, basic coding tasks, and straightforward multi-step instructions, DeepSeek V4 Pro is competent.

Humanity's Last Exam (HLE): top 10

On the hardest academic questions, DeepSeek V4 Pro ranks in the top 10 among AI models. This reflects DeepSeek's strong training on academic and scientific data — the model has broad knowledge across disciplines, even if it can't match the depth of reasoning of frontier models.


3. What does an IQ of 111 mean?

To put DeepSeek V4 Pro's IQ of 111 in context:

  • 85-115 (68% of people): Average intelligence
  • 115-130 (14%): Above average to high average
  • 130-145 (2%): Gifted — qualifies for Mensa (top 2%)
  • 145-160 (0.1%): Highly gifted / genius

DeepSeek V4 Pro, at 111, sits in the upper end of the average range — smarter than about 25% of humans, but not quite reaching the "above average" threshold of 115. If it were a person, it would be comparable to a bright college student or a capable professional — not a genius, but certainly competent.

This is notably lower than the frontier models:

  • Claude Opus 4.8: ~130 (gifted)
  • Gemini 3.1 Pro: 141 (highly gifted)
  • GPT-5.5: ~145 (genius)
  • Grok-4.20: 145 (genius)

The gap is significant — about 20-35 IQ points. But here's the key question: does the gap matter for your use case? For most everyday tasks — writing emails, summarizing documents, basic coding, translation — an IQ of 111 is more than sufficient. The frontier models' advantage shows up on the hardest problems: complex mathematical proofs, intricate multi-step reasoning, cutting-edge research questions.


4. Where DeepSeek V4 Pro excels

Cost-effectiveness

This is DeepSeek V4 Pro's defining feature. At 0.44/0.87 per 1M tokens, it's by far the cheapest major AI model. For high-volume applications — chatbots serving thousands of users, bulk document processing, large-scale data analysis — the cost savings are enormous. You could run 10-30 DeepSeek queries for the price of a single Claude or GPT-5.5 query.

Coding and technical tasks

DeepSeek V4 Pro is particularly strong at coding tasks. DeepSeek has invested heavily in code training data, and the model performs well on benchmarks like HumanEval and MBPP. For developers who need a coding assistant but can't afford frontier prices, DeepSeek V4 Pro is an excellent choice.

Open-source availability

DeepSeek V4 Pro's weights are available for local deployment. This means organizations can run the model on their own hardware, avoiding API costs entirely and keeping data private. For security-conscious organizations, researchers in restricted environments, and developers who need full control, this is a major advantage.

Multilingual support

DeepSeek V4 Pro is trained extensively on Chinese and English, with strong support for other major languages. For Chinese-language applications, DeepSeek is among the best models available — sometimes outperforming Western models on Chinese-language tasks.

Mathematical reasoning

While not at GPT-5.5's level, DeepSeek V4 Pro is competent at mathematical reasoning. It can solve most high-school and undergraduate-level math problems, and handles arithmetic and algebra reliably. For everyday mathematical tasks, it's more than adequate.

Speed

DeepSeek V4 Pro is fast — its smaller size and optimized architecture mean it generates responses quickly. For applications where latency matters — real-time chat, interactive tools, customer service — DeepSeek V4 Pro delivers responsive performance.


5. Where DeepSeek V4 Pro falls short

Raw IQ and complex reasoning

With an IQ of 111, DeepSeek V4 Pro is significantly below the frontier models on complex reasoning tasks. For the hardest problems — competition mathematics, advanced logic puzzles, multi-step abstract reasoning — the frontier models are substantially better. If your task requires genius-level reasoning, DeepSeek V4 Pro will struggle where Claude or GPT-5.5 would succeed.

Agentic tasks

On GDPval-AA v2, DeepSeek V4 Pro is in the top 10 but well behind Claude and GPT-5.5. For complex, multi-step real-world tasks — writing large software projects, managing research workflows, executing intricate plans — DeepSeek V4 Pro is less reliable. It can handle simple agentic tasks but struggles with complex ones.

Hallucination rate

DeepSeek V4 Pro has a moderate hallucination rate — higher than Claude (35.9%) and Gemini. It's more likely to confidently state incorrect information, particularly on niche or specialized topics. For applications where factual accuracy is critical (research, law, medicine), Claude is more reliable.

Emotional intelligence

DeepSeek V4 Pro's EQ (emotional intelligence) is estimated at ~105 — lower than all frontier models (Claude ~132, GPT-5.5 ~120, Gemini ~115, Grok ~110). Its responses tend to be more mechanical and less nuanced in understanding human emotions. For therapy apps, customer service, or sensitive human interactions, this is a limitation.

Safety and alignment

DeepSeek V4 Pro is subject to Chinese regulatory requirements, which means it may refuse to answer certain politically sensitive questions. This is different from the safety filters of Western models (which focus on harmful content) and may affect its usefulness for certain topics.


6. DeepSeek V4 Pro vs other models

Model Intelligence Index Mensa Norway IQ AI IQ est. Hallucination Cost/1M Open-source
DeepSeek V4 Pro 44.3 ~111 ~111 moderate 0.44/0.87 yes
Claude Opus 4.8 55.7 ~130 ~132 35.9% (lowest) 5/25 no
GPT-5.5 54.8 ~145 ~136 moderate 5/30 no
Gemini 3.1 Pro 46.5 141 ~131 low 2/12 no
Grok-4.20 top 5 145 ~140 moderate 3/15 no

The picture: DeepSeek V4 Pro is the budget champion — not the smartest, but by far the most affordable. Claude is the capable professional — best at real-world tasks and factual accuracy. GPT-5.5 is the test genius — best at math and visual reasoning. Gemini is the value champion — best cost-effectiveness among frontier models. Grok-4.20 is the raw IQ champion — highest IQ with a unique personality.

For users who need capable AI at the lowest possible price, DeepSeek V4 Pro is the clear choice. It won't match the frontier models on the hardest problems, but for 90% of everyday tasks, it's more than good enough — at a fraction of the cost.


7. What does this mean for humans?

DeepSeek V4 Pro operates at an IQ of ~111 — smarter than about 25% of humans. But:

  • IQ isn't everything: an IQ of 111 is above average and sufficient for most practical tasks
  • Cost matters more than IQ for most uses: for everyday applications, the 10x cost savings of DeepSeek often matter more than the 20-35 point IQ gap with frontier models
  • Open-source enables innovation: DeepSeek's open weights allow researchers and developers to build custom solutions, fine-tune for specific domains, and run locally
  • Knowledge ≠ understanding: like all AI models, DeepSeek has access to vast knowledge but doesn't truly "understand" the way a human does
  • No consciousness: high IQ doesn't mean sentience. DeepSeek is a sophisticated pattern-matching system, not a thinking being
  • Democratization of AI: DeepSeek V4 Pro proves that capable AI doesn't have to be expensive — it can be accessible to everyone

The most honest answer: DeepSeek V4 Pro is smarter than a quarter of humans on cognitive tests, and for most practical purposes, it's the smartest choice — not because it's the most intelligent, but because it delivers the most value per dollar.

How does your IQ compare to DeepSeek V4 Pro? Take NeuroLab's professional IQ assessment to find out.
Take the Test Now →


Conclusion

💡 Key takeaways
- DeepSeek V4 Pro's IQ is 111 on the Mensa Norway test — above average human, smarter than ~25% of people.

  • 44.3 on the Intelligence Index — top 10 globally, but ~10 points behind frontier models.
  • Lowest cost of any major model — 0.44/0.87 per 1M tokens, roughly 10x cheaper than Claude or GPT-5.5.
  • Open-source — weights available for local deployment, enabling privacy and customization.
  • Best for: cost-sensitive applications, coding, high-volume tasks, Chinese-language applications, local deployment.
  • Not best for: complex reasoning (frontier models win), agentic tasks (Claude wins), factual accuracy (Claude wins), emotional intelligence (Claude wins).
  • The lesson: intelligence isn't just about being the smartest — it's about being accessible. DeepSeek V4 Pro democratizes AI, bringing capable intelligence to everyone.