What is the IQ of Llama 4?
Dr. Daria Dudnik 10 min read 0 views

What is the IQ of Llama 4?

Llama 4 scores 112 on the Mensa Norway IQ test and ranks in the top 15 on the Intelligence Index. We break down every benchmark and explain what makes Meta's open-source model unique.

Llama 4: Meta's open-source champion

Llama 4 is the flagship open-source model from Meta AI, the research division of one of the world's largest technology companies. Released in early 2026, Llama 4 scores 112 on the Mensa Norway IQ test and ranks in the top 15 on the Artificial Analysis Intelligence Index v4.1. This places it above DeepSeek V4 Pro (111) but below Kimi K3 (118), Qwen 3.5 (115), and the frontier models (Claude ~130, Gemini 141, GPT-5.5 ~145, Grok-4.20 145).

What makes Llama 4 unique is its open-source nature at scale. Unlike closed models from OpenAI, Anthropic, or Google, Llama 4's weights are freely available for download and local deployment. This has made it the foundation for thousands of derivative models, fine-tunes, and applications — making Llama 4 not just a model, but an entire ecosystem.

Llama 4 is the people's model: a model that may not be the absolute smartest, but that anyone can download, modify, and run on their own hardware. For researchers, startups, and organizations that need full control over their AI infrastructure, Llama 4 offers something no closed model can: freedom.

Key scores:

  • Artificial Analysis Intelligence Index: top 15 (score ~43-44)
  • Mensa Norway IQ (TrackingAI): 112 (above average)
  • AI IQ composite (VexoWire): ~112 estimated
  • GDPval-AA v2: top 15
  • Humanity's Last Exam: top 15
  • Hallucination rate: moderate
  • Cost: free (self-hosted) or 0.20/0.60 per 1M tokens (via providers)
  • Parameters: 405B (largest variant)
  • Open-source: yes, weights freely available
  • Context: 128K tokens

1. The open-source revolution

Llama 4 represents Meta's commitment to open-source AI — the belief that AI should be accessible to everyone, not controlled by a few companies. While OpenAI, Anthropic, and Google keep their models behind APIs, Meta releases the full model weights, allowing anyone to download, study, modify, and deploy Llama 4 on their own hardware.

This has profound implications:

  • Democratization: anyone with sufficient hardware can run a state-of-the-art AI model
  • Customization: developers can fine-tune Llama 4 for specific domains, languages, or tasks
  • Privacy: organizations can run Llama 4 locally, ensuring data never leaves their infrastructure
  • Innovation: researchers can study the model's internals, leading to new breakthroughs
  • Ecosystem: thousands of derivative models build on Llama 4, creating a rich ecosystem
  • Cost: self-hosted Llama 4 is free (beyond hardware costs), making it the cheapest option at scale

The Llama ecosystem includes:

  • Llama 4 Base: the original pretrained model
  • Llama 4 Instruct: fine-tuned for instruction following
  • Llama 4 Guard: safety-filtered variant
  • Community fine-tunes: thousands of domain-specific variants
  • Quantized versions: compressed models that run on consumer hardware
  • Llama 4 Vision: multimodal variant with image understanding

No other AI model has spawned such a rich ecosystem. While GPT-5.5 and Claude may be smarter, they are black boxes. Llama 4 is transparent, modifiable, and community-driven.


2. Benchmark breakdown

Mensa Norway IQ test: 112 (above average)

The Mensa Norway IQ test consists of 36 visual pattern recognition questions. Llama 4 scores 112 — above the human average of 100, placing it in the "above average" range. This is comparable to DeepSeek V4 Pro (111), below Qwen 3.5 (115), Kimi K3 (118), and the frontier models (Claude ~130, Gemini 141, GPT-5.5 ~145, Grok-4.20 145).

An IQ of 112 means Llama 4 is smarter than approximately 79% of humans — the top 21% of human intelligence. If it were a person, it would be comparable to a capable university student or a skilled professional. Not a genius, but certainly bright and competent.

Artificial Analysis Intelligence Index: top 15

On the Intelligence Index, Llama 4 ranks in the top 15 globally with an estimated score of ~43-44. This places it competitive with DeepSeek V4 Pro (44.3) but behind Qwen 3.5 (~45-46), Kimi K3 (~45-47), Gemini (46.5), and the frontier models.

Breaking down the Index components:

  • GDPval-AA v2: Llama 4 performs adequately, benefiting from its large parameter count for diverse agentic tasks
  • AA-Omniscience: Llama 4 has a moderate hallucination rate — better than smaller models but worse than Claude
  • GPQA: Llama 4 performs reasonably on graduate-level science questions
  • HLE: Llama 4 ranks in the top 15 on Humanity's Last Exam
  • ARC-AGI: Llama 4 performs adequately on abstract reasoning
  • LMSYS Arena: Llama 4 gets solid human preference scores, particularly for open-ended tasks

AI IQ composite (VexoWire): ~112

The VexoWire composite IQ blends multiple tests (Mensa Norway, Mensa Denmark, Raven's, WAIS-IV). Llama 4's estimated composite IQ is ~112 — consistent with its Mensa Norway score. This places it in the "above average" range, smarter than about 79% of humans.

GDPval-AA v2: top 15

On the agentic benchmark, Llama 4 ranks in the top 15. Its large parameter count (405B) gives it substantial knowledge and reasoning capacity, but it trails the frontier models on complex multi-step tasks. For moderate agentic workloads, Llama 4 is competent.

Humanity's Last Exam (HLE): top 15

On the hardest academic questions, Llama 4 ranks in the top 15 among AI models. Meta's extensive training data and the model's large capacity help it perform well across disciplines, though it doesn't match the frontier models on the most specialized questions.


3. What does an IQ of 112 mean?

To put Llama 4's IQ of 112 in context:

  • 85-115 (68% of people): Average intelligence
  • 115-130 (14%): Above average to high average
  • 130-145 (2%): Gifted — qualifies for Mensa (top 2%)
  • 145-160 (0.1%): Highly gifted / genius

Llama 4, at 112, sits in the upper portion of the "average" range — smarter than about 79% of humans. If it were a person, it would be comparable to a capable university student or a skilled professional. Not a genius, but certainly bright and competent.

This is comparable to DeepSeek V4 Pro (111) but below the other models:

  • Qwen 3.5: 115 (above average)
  • Kimi K3: 118 (above average)
  • Claude Opus 4.8: ~130 (gifted)
  • Gemini 3.1 Pro: 141 (highly gifted)
  • GPT-5.5: ~145 (genius)
  • Grok-4.20: 145 (genius)

The gap to the frontier models is about 18-33 IQ points. But Llama 4 compensates with its open-source nature — for applications that require full control, privacy, or customization, Llama 4's openness can outweigh the IQ gap.


4. Where Llama 4 excels

Open-source freedom

Llama 4's weights are freely available for download and local deployment. This is its defining feature. No other top-tier model offers this level of openness. For organizations that need full control over their AI infrastructure, Llama 4 is the only option among top models.

Customization and fine-tuning

Because the weights are available, developers can fine-tune Llama 4 for specific domains, languages, or tasks. This has led to thousands of community fine-tunes — medical models, legal models, coding models, creative writing models, and more. No closed model offers this level of customization.

Privacy and data security

With Llama 4, organizations can run the model entirely on their own hardware, ensuring data never leaves their infrastructure. For healthcare, finance, defense, and other sensitive industries, this is critical. No data goes to third-party APIs.

Cost at scale

Self-hosted Llama 4 is free (beyond hardware costs). At scale, this can be dramatically cheaper than paying per-token API fees. For organizations with high inference volumes, the hardware investment pays for itself quickly.

Community ecosystem

The Llama ecosystem is unmatched. Thousands of developers contribute tools, fine-tunes, quantizations, and optimizations. This community-driven approach means Llama 4 benefits from collective innovation — improvements come from everywhere, not just from Meta.

Transparency

Unlike closed models, Llama 4's architecture and weights are transparent. Researchers can study how the model works, identify biases, understand failures, and propose improvements. This transparency is essential for scientific progress and responsible AI development.

Hardware optimization

The community has developed numerous quantized versions of Llama 4 that can run on consumer hardware — from 8-bit to 4-bit to 2-bit quantization. This means you can run a capable AI model on a high-end GPU, not just in a data center.


5. Where Llama 4 falls short

Raw IQ and complex reasoning

With an IQ of 112, Llama 4 is below the frontier models on complex reasoning tasks. For the hardest problems — competition mathematics, advanced logic puzzles, cutting-edge scientific reasoning — Claude, GPT-5.5, and Gemini are substantially better. The 18-33 IQ point gap is significant.

Hallucination rate

Llama 4 has a moderate hallucination rate — better than smaller models but worse than Claude (35.9%) and GPT-5.5. For applications where factual accuracy is critical — research, law, medicine — Claude remains more reliable.

Context window

With a 128K token context window, Llama 4 is adequate but not exceptional. Kimi K3 (2M), Gemini (1M), and Qwen 3.5 (256K) all offer larger context windows. For tasks that require processing very long documents, other models are better choices.

Multimodal capabilities

While Llama 4 Vision exists, its multimodal capabilities are less polished than Qwen 3.5 (text, image, audio, video) or Gemini (text, image, audio, video). For applications that require robust multimodal understanding, Qwen 3.5 or Gemini are better choices.

Emotional intelligence

Llama 4's EQ (emotional intelligence) is estimated at ~105 — lower than Claude (~132), GPT-5.5 (~120), Gemini (~115), and Qwen 3.5 (~107), but comparable to DeepSeek (~105). Its responses tend to be more functional and less emotionally nuanced. For therapy apps, customer service, or sensitive human interactions, Claude is the better choice.

Hardware requirements

The full 405B parameter model requires significant hardware to run — multiple high-end GPUs for inference. While quantized versions help, running Llama 4 at full quality is expensive in terms of hardware. For organizations without sufficient compute resources, API-based models may be more practical.

Safety and alignment

While Llama 4 Guard exists for safety filtering, the open-source nature means bad actors can remove safety measures. Meta provides guidelines but cannot enforce them. For applications requiring guaranteed safety, closed models with enforced guardrails may be more appropriate.


6. Llama 4 vs other models

Model Intelligence Index Mensa Norway IQ AI IQ est. Open-source Cost/1M Context Parameters
Llama 4 ~43-44 (top 15) 112 ~112 yes (free) free (self-host) 128K 405B
Claude Opus 4.8 55.7 ~130 ~132 no 5/25 200K unknown
GPT-5.5 54.8 ~145 ~136 no 5/30 400K unknown
Gemini 3.1 Pro 46.5 141 ~131 no 2/12 1M unknown
Grok-4.20 top 5 145 ~140 no 3/15 256K unknown
DeepSeek V4 Pro 44.3 ~111 ~111 yes 0.44/0.87 128K unknown
Qwen 3.5 ~45-46 (top 10) 115 ~115 yes 0.50/1.50 256K unknown
Kimi K3 ~45-47 (top 10) 118 ~118 no 0.55/2.19 2M unknown

The picture: Llama 4 is the open-source champion — not the smartest, but the most accessible. Claude is the capable professional — best at real-world tasks. GPT-5.5 is the test genius — best at math and visual reasoning. Gemini is the value champion — best cost-effectiveness among frontier models. Grok-4.20 is the raw IQ champion — highest IQ with personality. DeepSeek is the budget champion — cheapest API. Qwen 3.5 is the all-rounder — most versatile. And Kimi is the long-context champion — biggest context window.

For organizations that need open-source, privacy, customization, or cost-at-scale, Llama 4 is the clear choice. It's the people's model — not the smartest, but the most free.


7. What does this mean for humans?

Llama 4 operates at an IQ of ~112 — smarter than about 79% of humans. But:

  • IQ isn't everything: an IQ of 112 is above average and sufficient for most practical tasks
  • Openness matters more than raw IQ for many applications: for privacy-sensitive or customization-heavy applications, Llama 4's openness is more valuable than a higher IQ
  • Freedom enables innovation: Llama 4's open weights have spawned an entire ecosystem of innovation that closed models cannot match
  • Transparency builds trust: being able to inspect the model's internals builds trust that black-box models cannot
  • Knowledge ≠ understanding: like all AI models, Llama 4 has access to vast knowledge but doesn't truly "understand" the way a human does
  • No consciousness: high IQ doesn't mean sentience. Llama 4 is a sophisticated pattern-matching system, not a thinking being
  • Community drives progress: the Llama ecosystem proves that open collaboration can produce models competitive with billion-dollar closed systems

The most honest answer: Llama 4 is smarter than 79% of humans on cognitive tests, and for applications that require openness, privacy, or customization, it is the smartest choice — not because it's the most intelligent, but because it's the most free.

How does your IQ compare to Llama 4? Take NeuroLab's professional IQ assessment to find out.
Take the Test Now →


Conclusion

💡 Key takeaways
- Llama 4's IQ is 112 on the Mensa Norway test — above average, smarter than ~79% of humans.

  • Top 15 on the Intelligence Index — competitive with DeepSeek, behind Qwen 3.5, Kimi, and frontier models.
  • Fully open-source — weights freely available for download, study, and local deployment.
  • 405B parameters — one of the largest open-source models available.
  • Best for: privacy-sensitive applications, customization needs, cost-at-scale, research, self-hosted deployments, community-driven innovation.
  • Not best for: raw IQ (frontier models win), long context (Kimi wins), multimodal (Qwen 3.5 or Gemini win), emotional intelligence (Claude wins), guaranteed safety (closed models with enforced guardrails).
  • The lesson: intelligence isn't just about being the smartest — it's about being the most accessible. Llama 4 proves that open-source AI can compete with billion-dollar closed systems, democratizing access to advanced AI for everyone.