[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"$f321ifh7ge7j3e":3,"$f27qo82yht9ynw":104},{"success":4,"data":5},true,{"id":6,"slug":7,"coverImage":8,"author":9,"viewsCount":10,"createdAt":11,"title":12,"description":13,"lang":14,"contentHtml":15,"readingTime":16,"toc":17},"z8wcdr5vus4cvw8zr3p1ok1r","iq-llama-4","https:\u002F\u002Fimages.unsplash.com\u002Fphoto-1526374965328-4f18af581d06?w=1280","Dr. Daria Dudnik",0,"2026-08-01T10:00:01.170Z","What is the IQ of Llama 4?","Llama 4 scores 112 on the Mensa Norway IQ test and ranks in the top 15 on the Intelligence Index. We break down every benchmark and explain what makes Meta's open-source model unique.","en","\u003Ch2 id=\"llama-4-meta-s-open-source-champion\">Llama 4: Meta&#39;s open-source champion\u003C\u002Fh2>\n\u003Cp>Llama 4 is the flagship open-source model from Meta AI, the research division of one of the world&#39;s largest technology companies. Released in early 2026, Llama 4 scores \u003Cstrong>112\u003C\u002Fstrong> on the Mensa Norway IQ test and ranks in the top 15 on the Artificial Analysis Intelligence Index v4.1. This places it above DeepSeek V4 Pro (111) but below Kimi K3 (118), Qwen 3.5 (115), and the frontier models (Claude ~130, Gemini 141, GPT-5.5 ~145, Grok-4.20 145).\u003C\u002Fp>\n\u003Cp>What makes Llama 4 unique is its \u003Cstrong>open-source nature at scale\u003C\u002Fstrong>. Unlike closed models from OpenAI, Anthropic, or Google, Llama 4&#39;s weights are freely available for download and local deployment. This has made it the foundation for thousands of derivative models, fine-tunes, and applications — making Llama 4 not just a model, but an entire ecosystem.\u003C\u002Fp>\n\u003Cp>Llama 4 is the \u003Cstrong>people&#39;s model\u003C\u002Fstrong>: a model that may not be the absolute smartest, but that anyone can download, modify, and run on their own hardware. For researchers, startups, and organizations that need full control over their AI infrastructure, Llama 4 offers something no closed model can: freedom.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>Key scores:\u003C\u002Fstrong>\u003C\u002Fp>\n\u003Cul>\n\u003Cli>\u003Cstrong>Artificial Analysis Intelligence Index\u003C\u002Fstrong>: top 15 (score ~43-44)\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Mensa Norway IQ (TrackingAI)\u003C\u002Fstrong>: 112 (above average)\u003C\u002Fli>\n\u003Cli>\u003Cstrong>AI IQ composite (VexoWire)\u003C\u002Fstrong>: ~112 estimated\u003C\u002Fli>\n\u003Cli>\u003Cstrong>GDPval-AA v2\u003C\u002Fstrong>: top 15\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Humanity&#39;s Last Exam\u003C\u002Fstrong>: top 15\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Hallucination rate\u003C\u002Fstrong>: moderate\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Cost\u003C\u002Fstrong>: free (self-hosted) or 0.20\u002F0.60 per 1M tokens (via providers)\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Parameters\u003C\u002Fstrong>: 405B (largest variant)\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Open-source\u003C\u002Fstrong>: yes, weights freely available\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Context\u003C\u002Fstrong>: 128K tokens\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Chr>\n\u003Ch2 id=\"1-the-open-source-revolution\">1. The open-source revolution\u003C\u002Fh2>\n\u003Cp>Llama 4 represents Meta&#39;s commitment to open-source AI — the belief that AI should be accessible to everyone, not controlled by a few companies. While OpenAI, Anthropic, and Google keep their models behind APIs, Meta releases the full model weights, allowing anyone to download, study, modify, and deploy Llama 4 on their own hardware.\u003C\u002Fp>\n\u003Cp>This has profound implications:\u003C\u002Fp>\n\u003Cul>\n\u003Cli>\u003Cstrong>Democratization\u003C\u002Fstrong>: anyone with sufficient hardware can run a state-of-the-art AI model\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Customization\u003C\u002Fstrong>: developers can fine-tune Llama 4 for specific domains, languages, or tasks\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Privacy\u003C\u002Fstrong>: organizations can run Llama 4 locally, ensuring data never leaves their infrastructure\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Innovation\u003C\u002Fstrong>: researchers can study the model&#39;s internals, leading to new breakthroughs\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Ecosystem\u003C\u002Fstrong>: thousands of derivative models build on Llama 4, creating a rich ecosystem\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Cost\u003C\u002Fstrong>: self-hosted Llama 4 is free (beyond hardware costs), making it the cheapest option at scale\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Cp>The Llama ecosystem includes:\u003C\u002Fp>\n\u003Cul>\n\u003Cli>\u003Cstrong>Llama 4 Base\u003C\u002Fstrong>: the original pretrained model\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Llama 4 Instruct\u003C\u002Fstrong>: fine-tuned for instruction following\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Llama 4 Guard\u003C\u002Fstrong>: safety-filtered variant\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Community fine-tunes\u003C\u002Fstrong>: thousands of domain-specific variants\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Quantized versions\u003C\u002Fstrong>: compressed models that run on consumer hardware\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Llama 4 Vision\u003C\u002Fstrong>: multimodal variant with image understanding\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Cp>No other AI model has spawned such a rich ecosystem. While GPT-5.5 and Claude may be smarter, they are black boxes. Llama 4 is transparent, modifiable, and community-driven.\u003C\u002Fp>\n\u003Chr>\n\u003Ch2 id=\"2-benchmark-breakdown\">2. Benchmark breakdown\u003C\u002Fh2>\n\u003Ch3 id=\"mensa-norway-iq-test-112-above-average\">Mensa Norway IQ test: 112 (above average)\u003C\u002Fh3>\n\u003Cp>The Mensa Norway IQ test consists of 36 visual pattern recognition questions. Llama 4 scores \u003Cstrong>112\u003C\u002Fstrong> — above the human average of 100, placing it in the &quot;above average&quot; range. This is comparable to DeepSeek V4 Pro (111), below Qwen 3.5 (115), Kimi K3 (118), and the frontier models (Claude ~130, Gemini 141, GPT-5.5 ~145, Grok-4.20 145).\u003C\u002Fp>\n\u003Cp>An IQ of 112 means Llama 4 is smarter than approximately 79% of humans — the top 21% of human intelligence. If it were a person, it would be comparable to a capable university student or a skilled professional. Not a genius, but certainly bright and competent.\u003C\u002Fp>\n\u003Ch3 id=\"artificial-analysis-intelligence-index-top-15\">Artificial Analysis Intelligence Index: top 15\u003C\u002Fh3>\n\u003Cp>On the Intelligence Index, Llama 4 ranks in the top 15 globally with an estimated score of \u003Cstrong>~43-44\u003C\u002Fstrong>. This places it competitive with DeepSeek V4 Pro (44.3) but behind Qwen 3.5 (~45-46), Kimi K3 (~45-47), Gemini (46.5), and the frontier models.\u003C\u002Fp>\n\u003Cp>Breaking down the Index components:\u003C\u002Fp>\n\u003Cul>\n\u003Cli>\u003Cstrong>GDPval-AA v2\u003C\u002Fstrong>: Llama 4 performs adequately, benefiting from its large parameter count for diverse agentic tasks\u003C\u002Fli>\n\u003Cli>\u003Cstrong>AA-Omniscience\u003C\u002Fstrong>: Llama 4 has a moderate hallucination rate — better than smaller models but worse than Claude\u003C\u002Fli>\n\u003Cli>\u003Cstrong>GPQA\u003C\u002Fstrong>: Llama 4 performs reasonably on graduate-level science questions\u003C\u002Fli>\n\u003Cli>\u003Cstrong>HLE\u003C\u002Fstrong>: Llama 4 ranks in the top 15 on Humanity&#39;s Last Exam\u003C\u002Fli>\n\u003Cli>\u003Cstrong>ARC-AGI\u003C\u002Fstrong>: Llama 4 performs adequately on abstract reasoning\u003C\u002Fli>\n\u003Cli>\u003Cstrong>LMSYS Arena\u003C\u002Fstrong>: Llama 4 gets solid human preference scores, particularly for open-ended tasks\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Ch3 id=\"ai-iq-composite-vexowire-112\">AI IQ composite (VexoWire): ~112\u003C\u002Fh3>\n\u003Cp>The VexoWire composite IQ blends multiple tests (Mensa Norway, Mensa Denmark, Raven&#39;s, WAIS-IV). Llama 4&#39;s estimated composite IQ is \u003Cstrong>~112\u003C\u002Fstrong> — consistent with its Mensa Norway score. This places it in the &quot;above average&quot; range, smarter than about 79% of humans.\u003C\u002Fp>\n\u003Ch3 id=\"gdpval-aa-v2-top-15\">GDPval-AA v2: top 15\u003C\u002Fh3>\n\u003Cp>On the agentic benchmark, Llama 4 ranks in the top 15. Its large parameter count (405B) gives it substantial knowledge and reasoning capacity, but it trails the frontier models on complex multi-step tasks. For moderate agentic workloads, Llama 4 is competent.\u003C\u002Fp>\n\u003Ch3 id=\"humanity-s-last-exam-hle-top-15\">Humanity&#39;s Last Exam (HLE): top 15\u003C\u002Fh3>\n\u003Cp>On the hardest academic questions, Llama 4 ranks in the top 15 among AI models. Meta&#39;s extensive training data and the model&#39;s large capacity help it perform well across disciplines, though it doesn&#39;t match the frontier models on the most specialized questions.\u003C\u002Fp>\n\u003Chr>\n\u003Ch2 id=\"3-what-does-an-iq-of-112-mean\">3. What does an IQ of 112 mean?\u003C\u002Fh2>\n\u003Cp>To put Llama 4&#39;s IQ of 112 in context:\u003C\u002Fp>\n\u003Cul>\n\u003Cli>\u003Cstrong>85-115 (68% of people)\u003C\u002Fstrong>: Average intelligence\u003C\u002Fli>\n\u003Cli>\u003Cstrong>115-130 (14%)\u003C\u002Fstrong>: Above average to high average\u003C\u002Fli>\n\u003Cli>\u003Cstrong>130-145 (2%)\u003C\u002Fstrong>: Gifted — qualifies for Mensa (top 2%)\u003C\u002Fli>\n\u003Cli>\u003Cstrong>145-160 (0.1%)\u003C\u002Fstrong>: Highly gifted \u002F genius\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Cp>Llama 4, at 112, sits in the \u003Cstrong>upper portion of the &quot;average&quot; range\u003C\u002Fstrong> — smarter than about 79% of humans. If it were a person, it would be comparable to a capable university student or a skilled professional. Not a genius, but certainly bright and competent.\u003C\u002Fp>\n\u003Cp>This is comparable to DeepSeek V4 Pro (111) but below the other models:\u003C\u002Fp>\n\u003Cul>\n\u003Cli>Qwen 3.5: 115 (above average)\u003C\u002Fli>\n\u003Cli>Kimi K3: 118 (above average)\u003C\u002Fli>\n\u003Cli>Claude Opus 4.8: ~130 (gifted)\u003C\u002Fli>\n\u003Cli>Gemini 3.1 Pro: 141 (highly gifted)\u003C\u002Fli>\n\u003Cli>GPT-5.5: ~145 (genius)\u003C\u002Fli>\n\u003Cli>Grok-4.20: 145 (genius)\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Cp>The gap to the frontier models is about 18-33 IQ points. But Llama 4 compensates with its open-source nature — for applications that require full control, privacy, or customization, Llama 4&#39;s openness can outweigh the IQ gap.\u003C\u002Fp>\n\u003Chr>\n\u003Ch2 id=\"4-where-llama-4-excels\">4. Where Llama 4 excels\u003C\u002Fh2>\n\u003Ch3 id=\"open-source-freedom\">Open-source freedom\u003C\u002Fh3>\n\u003Cp>Llama 4&#39;s weights are \u003Cstrong>freely available\u003C\u002Fstrong> for download and local deployment. This is its defining feature. No other top-tier model offers this level of openness. For organizations that need full control over their AI infrastructure, Llama 4 is the only option among top models.\u003C\u002Fp>\n\u003Ch3 id=\"customization-and-fine-tuning\">Customization and fine-tuning\u003C\u002Fh3>\n\u003Cp>Because the weights are available, developers can \u003Cstrong>fine-tune Llama 4\u003C\u002Fstrong> for specific domains, languages, or tasks. This has led to thousands of community fine-tunes — medical models, legal models, coding models, creative writing models, and more. No closed model offers this level of customization.\u003C\u002Fp>\n\u003Ch3 id=\"privacy-and-data-security\">Privacy and data security\u003C\u002Fh3>\n\u003Cp>With Llama 4, organizations can run the model \u003Cstrong>entirely on their own hardware\u003C\u002Fstrong>, ensuring data never leaves their infrastructure. For healthcare, finance, defense, and other sensitive industries, this is critical. No data goes to third-party APIs.\u003C\u002Fp>\n\u003Ch3 id=\"cost-at-scale\">Cost at scale\u003C\u002Fh3>\n\u003Cp>Self-hosted Llama 4 is \u003Cstrong>free\u003C\u002Fstrong> (beyond hardware costs). At scale, this can be dramatically cheaper than paying per-token API fees. For organizations with high inference volumes, the hardware investment pays for itself quickly.\u003C\u002Fp>\n\u003Ch3 id=\"community-ecosystem\">Community ecosystem\u003C\u002Fh3>\n\u003Cp>The Llama ecosystem is unmatched. Thousands of developers contribute tools, fine-tunes, quantizations, and optimizations. This community-driven approach means Llama 4 benefits from collective innovation — improvements come from everywhere, not just from Meta.\u003C\u002Fp>\n\u003Ch3 id=\"transparency\">Transparency\u003C\u002Fh3>\n\u003Cp>Unlike closed models, Llama 4&#39;s architecture and weights are \u003Cstrong>transparent\u003C\u002Fstrong>. Researchers can study how the model works, identify biases, understand failures, and propose improvements. This transparency is essential for scientific progress and responsible AI development.\u003C\u002Fp>\n\u003Ch3 id=\"hardware-optimization\">Hardware optimization\u003C\u002Fh3>\n\u003Cp>The community has developed numerous \u003Cstrong>quantized versions\u003C\u002Fstrong> of Llama 4 that can run on consumer hardware — from 8-bit to 4-bit to 2-bit quantization. This means you can run a capable AI model on a high-end GPU, not just in a data center.\u003C\u002Fp>\n\u003Chr>\n\u003Ch2 id=\"5-where-llama-4-falls-short\">5. Where Llama 4 falls short\u003C\u002Fh2>\n\u003Ch3 id=\"raw-iq-and-complex-reasoning\">Raw IQ and complex reasoning\u003C\u002Fh3>\n\u003Cp>With an IQ of 112, Llama 4 is below the frontier models on complex reasoning tasks. For the hardest problems — competition mathematics, advanced logic puzzles, cutting-edge scientific reasoning — Claude, GPT-5.5, and Gemini are substantially better. The 18-33 IQ point gap is significant.\u003C\u002Fp>\n\u003Ch3 id=\"hallucination-rate\">Hallucination rate\u003C\u002Fh3>\n\u003Cp>Llama 4 has a \u003Cstrong>moderate hallucination rate\u003C\u002Fstrong> — better than smaller models but worse than Claude (35.9%) and GPT-5.5. For applications where factual accuracy is critical — research, law, medicine — Claude remains more reliable.\u003C\u002Fp>\n\u003Ch3 id=\"context-window\">Context window\u003C\u002Fh3>\n\u003Cp>With a 128K token context window, Llama 4 is adequate but not exceptional. Kimi K3 (2M), Gemini (1M), and Qwen 3.5 (256K) all offer larger context windows. For tasks that require processing very long documents, other models are better choices.\u003C\u002Fp>\n\u003Ch3 id=\"multimodal-capabilities\">Multimodal capabilities\u003C\u002Fh3>\n\u003Cp>While Llama 4 Vision exists, its multimodal capabilities are less polished than Qwen 3.5 (text, image, audio, video) or Gemini (text, image, audio, video). For applications that require robust multimodal understanding, Qwen 3.5 or Gemini are better choices.\u003C\u002Fp>\n\u003Ch3 id=\"emotional-intelligence\">Emotional intelligence\u003C\u002Fh3>\n\u003Cp>Llama 4&#39;s EQ (emotional intelligence) is estimated at ~105 — lower than Claude (~132), GPT-5.5 (~120), Gemini (~115), and Qwen 3.5 (~107), but comparable to DeepSeek (~105). Its responses tend to be more functional and less emotionally nuanced. For therapy apps, customer service, or sensitive human interactions, Claude is the better choice.\u003C\u002Fp>\n\u003Ch3 id=\"hardware-requirements\">Hardware requirements\u003C\u002Fh3>\n\u003Cp>The full 405B parameter model requires \u003Cstrong>significant hardware\u003C\u002Fstrong> to run — multiple high-end GPUs for inference. While quantized versions help, running Llama 4 at full quality is expensive in terms of hardware. For organizations without sufficient compute resources, API-based models may be more practical.\u003C\u002Fp>\n\u003Ch3 id=\"safety-and-alignment\">Safety and alignment\u003C\u002Fh3>\n\u003Cp>While Llama 4 Guard exists for safety filtering, the open-source nature means bad actors can remove safety measures. Meta provides guidelines but cannot enforce them. For applications requiring guaranteed safety, closed models with enforced guardrails may be more appropriate.\u003C\u002Fp>\n\u003Chr>\n\u003Ch2 id=\"6-llama-4-vs-other-models\">6. Llama 4 vs other models\u003C\u002Fh2>\n\u003Cdiv class=\"article-table-wrapper\">\u003Ctable>\n\u003Cthead>\n\u003Ctr>\n\u003Cth>Model\u003C\u002Fth>\n\u003Cth>Intelligence Index\u003C\u002Fth>\n\u003Cth>Mensa Norway IQ\u003C\u002Fth>\n\u003Cth>AI IQ est.\u003C\u002Fth>\n\u003Cth>Open-source\u003C\u002Fth>\n\u003Cth>Cost\u002F1M\u003C\u002Fth>\n\u003Cth>Context\u003C\u002Fth>\n\u003Cth>Parameters\u003C\u002Fth>\n\u003C\u002Ftr>\n\u003C\u002Fthead>\n\u003Ctbody>\u003Ctr>\n\u003Ctd>Llama 4\u003C\u002Ftd>\n\u003Ctd>~43-44 (top 15)\u003C\u002Ftd>\n\u003Ctd>112\u003C\u002Ftd>\n\u003Ctd>~112\u003C\u002Ftd>\n\u003Ctd>yes (free)\u003C\u002Ftd>\n\u003Ctd>free (self-host)\u003C\u002Ftd>\n\u003Ctd>128K\u003C\u002Ftd>\n\u003Ctd>405B\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003Ctr>\n\u003Ctd>Claude Opus 4.8\u003C\u002Ftd>\n\u003Ctd>55.7\u003C\u002Ftd>\n\u003Ctd>~130\u003C\u002Ftd>\n\u003Ctd>~132\u003C\u002Ftd>\n\u003Ctd>no\u003C\u002Ftd>\n\u003Ctd>5\u002F25\u003C\u002Ftd>\n\u003Ctd>200K\u003C\u002Ftd>\n\u003Ctd>unknown\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003Ctr>\n\u003Ctd>GPT-5.5\u003C\u002Ftd>\n\u003Ctd>54.8\u003C\u002Ftd>\n\u003Ctd>~145\u003C\u002Ftd>\n\u003Ctd>~136\u003C\u002Ftd>\n\u003Ctd>no\u003C\u002Ftd>\n\u003Ctd>5\u002F30\u003C\u002Ftd>\n\u003Ctd>400K\u003C\u002Ftd>\n\u003Ctd>unknown\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003Ctr>\n\u003Ctd>Gemini 3.1 Pro\u003C\u002Ftd>\n\u003Ctd>46.5\u003C\u002Ftd>\n\u003Ctd>141\u003C\u002Ftd>\n\u003Ctd>~131\u003C\u002Ftd>\n\u003Ctd>no\u003C\u002Ftd>\n\u003Ctd>2\u002F12\u003C\u002Ftd>\n\u003Ctd>1M\u003C\u002Ftd>\n\u003Ctd>unknown\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003Ctr>\n\u003Ctd>Grok-4.20\u003C\u002Ftd>\n\u003Ctd>top 5\u003C\u002Ftd>\n\u003Ctd>145\u003C\u002Ftd>\n\u003Ctd>~140\u003C\u002Ftd>\n\u003Ctd>no\u003C\u002Ftd>\n\u003Ctd>3\u002F15\u003C\u002Ftd>\n\u003Ctd>256K\u003C\u002Ftd>\n\u003Ctd>unknown\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003Ctr>\n\u003Ctd>DeepSeek V4 Pro\u003C\u002Ftd>\n\u003Ctd>44.3\u003C\u002Ftd>\n\u003Ctd>~111\u003C\u002Ftd>\n\u003Ctd>~111\u003C\u002Ftd>\n\u003Ctd>yes\u003C\u002Ftd>\n\u003Ctd>0.44\u002F0.87\u003C\u002Ftd>\n\u003Ctd>128K\u003C\u002Ftd>\n\u003Ctd>unknown\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003Ctr>\n\u003Ctd>Qwen 3.5\u003C\u002Ftd>\n\u003Ctd>~45-46 (top 10)\u003C\u002Ftd>\n\u003Ctd>115\u003C\u002Ftd>\n\u003Ctd>~115\u003C\u002Ftd>\n\u003Ctd>yes\u003C\u002Ftd>\n\u003Ctd>0.50\u002F1.50\u003C\u002Ftd>\n\u003Ctd>256K\u003C\u002Ftd>\n\u003Ctd>unknown\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003Ctr>\n\u003Ctd>Kimi K3\u003C\u002Ftd>\n\u003Ctd>~45-47 (top 10)\u003C\u002Ftd>\n\u003Ctd>118\u003C\u002Ftd>\n\u003Ctd>~118\u003C\u002Ftd>\n\u003Ctd>no\u003C\u002Ftd>\n\u003Ctd>0.55\u002F2.19\u003C\u002Ftd>\n\u003Ctd>2M\u003C\u002Ftd>\n\u003Ctd>unknown\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003C\u002Ftbody>\u003C\u002Ftable>\u003C\u002Fdiv>\n\u003Cp>The picture: Llama 4 is the \u003Cstrong>open-source champion\u003C\u002Fstrong> — not the smartest, but the most accessible. Claude is the \u003Cstrong>capable professional\u003C\u002Fstrong> — best at real-world tasks. GPT-5.5 is the \u003Cstrong>test genius\u003C\u002Fstrong> — best at math and visual reasoning. Gemini is the \u003Cstrong>value champion\u003C\u002Fstrong> — best cost-effectiveness among frontier models. Grok-4.20 is the \u003Cstrong>raw IQ champion\u003C\u002Fstrong> — highest IQ with personality. DeepSeek is the \u003Cstrong>budget champion\u003C\u002Fstrong> — cheapest API. Qwen 3.5 is the \u003Cstrong>all-rounder\u003C\u002Fstrong> — most versatile. And Kimi is the \u003Cstrong>long-context champion\u003C\u002Fstrong> — biggest context window.\u003C\u002Fp>\n\u003Cp>For organizations that need open-source, privacy, customization, or cost-at-scale, Llama 4 is the clear choice. It&#39;s the people&#39;s model — not the smartest, but the most free.\u003C\u002Fp>\n\u003Chr>\n\u003Ch2 id=\"7-what-does-this-mean-for-humans\">7. What does this mean for humans?\u003C\u002Fh2>\n\u003Cp>Llama 4 operates at an IQ of ~112 — smarter than about 79% of humans. But:\u003C\u002Fp>\n\u003Cul>\n\u003Cli>\u003Cstrong>IQ isn&#39;t everything\u003C\u002Fstrong>: an IQ of 112 is above average and sufficient for most practical tasks\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Openness matters more than raw IQ for many applications\u003C\u002Fstrong>: for privacy-sensitive or customization-heavy applications, Llama 4&#39;s openness is more valuable than a higher IQ\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Freedom enables innovation\u003C\u002Fstrong>: Llama 4&#39;s open weights have spawned an entire ecosystem of innovation that closed models cannot match\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Transparency builds trust\u003C\u002Fstrong>: being able to inspect the model&#39;s internals builds trust that black-box models cannot\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Knowledge ≠ understanding\u003C\u002Fstrong>: like all AI models, Llama 4 has access to vast knowledge but doesn&#39;t truly &quot;understand&quot; the way a human does\u003C\u002Fli>\n\u003Cli>\u003Cstrong>No consciousness\u003C\u002Fstrong>: high IQ doesn&#39;t mean sentience. Llama 4 is a sophisticated pattern-matching system, not a thinking being\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Community drives progress\u003C\u002Fstrong>: the Llama ecosystem proves that open collaboration can produce models competitive with billion-dollar closed systems\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Cp>The most honest answer: Llama 4 is \u003Cstrong>smarter than 79% of humans on cognitive tests\u003C\u002Fstrong>, and for applications that require openness, privacy, or customization, it is the smartest choice — not because it&#39;s the most intelligent, but because it&#39;s the most free.\u003C\u002Fp>\n\u003Cp>\n    \u003Cdiv class=\"article-quiz-cta\">\n      \u003Cdiv class=\"article-quiz-cta__text\">How does your IQ compare to Llama 4? Take NeuroLab&#39;s professional IQ assessment to find out.\u003C\u002Fdiv>\n      \u003Ca href=\"\u002Fen\u002Ftests\u002Fiq-standard\" class=\"article-quiz-cta__button\">\n        Take the Test Now →\n      \u003C\u002Fa>\n    \u003C\u002Fdiv>\u003C\u002Fp>\n\u003Chr>\n\u003Ch2 id=\"conclusion\">Conclusion\u003C\u002Fh2>\n\u003Cp>\n    \u003Cdiv class=\"article-callout\">\n      \u003Cdiv class=\"article-callout__title\">\n        \u003Cspan class=\"article-callout__icon\">💡\u003C\u002Fspan>\n        Key takeaways\n      \u003C\u002Fdiv>\n      \u003Cdiv class=\"article-callout__body\">- \u003Cstrong>Llama 4&#39;s IQ is 112\u003C\u002Fstrong> on the Mensa Norway test — above average, smarter than ~79% of humans.\u003C\u002Fp>\n\u003Cul>\n\u003Cli>\u003Cstrong>Top 15 on the Intelligence Index\u003C\u002Fstrong> — competitive with DeepSeek, behind Qwen 3.5, Kimi, and frontier models.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Fully open-source\u003C\u002Fstrong> — weights freely available for download, study, and local deployment.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>405B parameters\u003C\u002Fstrong> — one of the largest open-source models available.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Best for\u003C\u002Fstrong>: privacy-sensitive applications, customization needs, cost-at-scale, research, self-hosted deployments, community-driven innovation.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Not best for\u003C\u002Fstrong>: raw IQ (frontier models win), long context (Kimi wins), multimodal (Qwen 3.5 or Gemini win), emotional intelligence (Claude wins), guaranteed safety (closed models with enforced guardrails).\u003C\u002Fli>\n\u003Cli>\u003Cstrong>The lesson\u003C\u002Fstrong>: intelligence isn&#39;t just about being the smartest — it&#39;s about being the most accessible. Llama 4 proves that open-source AI can compete with billion-dollar closed systems, democratizing access to advanced AI for everyone.\u003C\u002Fdiv>\n    \u003C\u002Fdiv>\u003C\u002Fli>\n\u003C\u002Ful>\n",10,[18,22,25,28,32,35,38,41,44,47,50,53,56,59,62,65,68,71,74,77,80,83,86,89,92,95,98,101],{"id":19,"text":20,"level":21},"llama-4-meta-s-open-source-champion","Llama 4: Meta's open-source champion",2,{"id":23,"text":24,"level":21},"1-the-open-source-revolution","1. The open-source revolution",{"id":26,"text":27,"level":21},"2-benchmark-breakdown","2. Benchmark breakdown",{"id":29,"text":30,"level":31},"mensa-norway-iq-test-112-above-average","Mensa Norway IQ test: 112 (above average)",3,{"id":33,"text":34,"level":31},"artificial-analysis-intelligence-index-top-15","Artificial Analysis Intelligence Index: top 15",{"id":36,"text":37,"level":31},"ai-iq-composite-vexowire-112","AI IQ composite (VexoWire): ~112",{"id":39,"text":40,"level":31},"gdpval-aa-v2-top-15","GDPval-AA v2: top 15",{"id":42,"text":43,"level":31},"humanity-s-last-exam-hle-top-15","Humanity's Last Exam (HLE): top 15",{"id":45,"text":46,"level":21},"3-what-does-an-iq-of-112-mean","3. What does an IQ of 112 mean?",{"id":48,"text":49,"level":21},"4-where-llama-4-excels","4. Where Llama 4 excels",{"id":51,"text":52,"level":31},"open-source-freedom","Open-source freedom",{"id":54,"text":55,"level":31},"customization-and-fine-tuning","Customization and fine-tuning",{"id":57,"text":58,"level":31},"privacy-and-data-security","Privacy and data security",{"id":60,"text":61,"level":31},"cost-at-scale","Cost at scale",{"id":63,"text":64,"level":31},"community-ecosystem","Community ecosystem",{"id":66,"text":67,"level":31},"transparency","Transparency",{"id":69,"text":70,"level":31},"hardware-optimization","Hardware optimization",{"id":72,"text":73,"level":21},"5-where-llama-4-falls-short","5. Where Llama 4 falls short",{"id":75,"text":76,"level":31},"raw-iq-and-complex-reasoning","Raw IQ and complex reasoning",{"id":78,"text":79,"level":31},"hallucination-rate","Hallucination rate",{"id":81,"text":82,"level":31},"context-window","Context window",{"id":84,"text":85,"level":31},"multimodal-capabilities","Multimodal capabilities",{"id":87,"text":88,"level":31},"emotional-intelligence","Emotional intelligence",{"id":90,"text":91,"level":31},"hardware-requirements","Hardware requirements",{"id":93,"text":94,"level":31},"safety-and-alignment","Safety and alignment",{"id":96,"text":97,"level":21},"6-llama-4-vs-other-models","6. Llama 4 vs other models",{"id":99,"text":100,"level":21},"7-what-does-this-mean-for-humans","7. What does this mean for humans?",{"id":102,"text":103,"level":21},"conclusion","Conclusion",{"success":4,"data":105,"pagination":274},[106,114,121,129,130,138,145,152,160,167,174,181,188,195,203,210,217,224,232,239,246,253,260,267],{"id":107,"slug":108,"coverImage":109,"author":9,"viewsCount":10,"createdAt":110,"title":111,"description":112,"lang":14,"readingTime":113},"jy0egbpgrn30am622jy7xfrw","adhd-and-iq","https:\u002F\u002Fimages.unsplash.com\u002Fphoto-1507812984078-917231447e3e?w=1280","2026-08-01T10:43:58.026Z","ADHD and IQ: When Attention Deficit Hides High Intelligence","Can ADHD mask high IQ? Yes. Research shows that attention deficits can artificially lower IQ test scores by 10-15 points — hiding gifted minds behind poor focus.",7,{"id":115,"slug":116,"coverImage":117,"author":9,"viewsCount":10,"createdAt":118,"title":119,"description":120,"lang":14,"readingTime":113},"dvlwozf321i8jlk9sfl390k1","online-vs-certified-iq-tests","https:\u002F\u002Fimages.unsplash.com\u002Fphoto-1499209974431-9fccce68dc04?w=1280","2026-08-01T10:42:47.143Z","Online IQ Tests vs Certified Assessments: How Accurate Are They?","You scored 140 on a free online IQ test. But what does that actually mean? We compare online tests with certified assessments — and the results may surprise you.",{"id":122,"slug":123,"coverImage":124,"author":9,"viewsCount":10,"createdAt":125,"title":126,"description":127,"lang":14,"readingTime":128},"maxnlmh4wqwxjuidbe2cb1jh","iq-above-160","https:\u002F\u002Fimages.unsplash.com\u002Fphoto-1454165804606-c3d57bc86b40?w=1280","2026-08-01T10:41:43.616Z","IQ Above 160: What Life Looks Like at the Extreme","Only 1 in 30,000 people have an IQ above 160. We explore what life is like at the extreme edge of human intelligence — the gifts, the struggles, and the myths.",8,{"id":6,"slug":7,"coverImage":8,"author":9,"viewsCount":10,"createdAt":11,"title":12,"description":13,"lang":14,"readingTime":16},{"id":131,"slug":132,"coverImage":133,"author":9,"viewsCount":134,"createdAt":135,"title":136,"description":137,"lang":14,"readingTime":16},"pxehdebhlpxfs0h5qm2m23e9","iq-qwen-35","https:\u002F\u002Fimages.unsplash.com\u002Fphoto-1518186285589-2f7649de83e0?w=1280",1,"2026-08-01T09:38:28.903Z","What is the IQ of Qwen 3.5?","Qwen 3.5 scores 115 on the Mensa Norway IQ test and ranks in the top 10 on the Intelligence Index. We break down every benchmark and explain what makes Alibaba's model unique.",{"id":139,"slug":140,"coverImage":141,"author":9,"viewsCount":21,"createdAt":142,"title":143,"description":144,"lang":14,"readingTime":16},"st1nxd9ef5blcslk68p83pj1","iq-kimi-k3","https:\u002F\u002Fimages.unsplash.com\u002Fphoto-1620712943524-bccbe725b837?w=1280","2026-08-01T09:29:37.632Z","What is the IQ of Kimi K3?","Kimi K3 scores 118 on the Mensa Norway IQ test and ranks in the top 10 on the Intelligence Index. We break down every benchmark and explain what makes Moonshot AI's model unique.",{"id":146,"slug":147,"coverImage":148,"author":9,"viewsCount":10,"createdAt":149,"title":150,"description":151,"lang":14,"readingTime":16},"u65iu8mijuzx0t4g1981phue","iq-deepseek-v4","https:\u002F\u002Fimages.unsplash.com\u002Fphoto-1677372139014-394e5b51f79b?w=1280","2026-08-01T09:20:10.570Z","What is the IQ of DeepSeek V4?","DeepSeek V4 Pro scores 111 on the Mensa Norway IQ test and 44.3 on the Intelligence Index. We break down every benchmark and explain why it's the best budget AI model.",{"id":153,"slug":154,"coverImage":155,"author":9,"viewsCount":156,"createdAt":157,"title":158,"description":159,"lang":14,"readingTime":16},"lci1u8drubje01qwq1vah01m","iq-grok-420","https:\u002F\u002Fimages.unsplash.com\u002Fphoto-1635776062127-d379b0ba009b?w=1280",6,"2026-07-31T20:28:41.558Z","What is the IQ of Grok-4.20?","Grok-4.20 scores 145 on the Mensa Norway IQ test — tied for #1 among all AI models. We break down every benchmark and explain what makes xAI's model unique.",{"id":161,"slug":162,"coverImage":163,"author":9,"viewsCount":21,"createdAt":164,"title":165,"description":166,"lang":14,"readingTime":16},"n32qvrqd1er095jgz98qg73l","iq-gemini-31-pro","https:\u002F\u002Fimages.unsplash.com\u002Fphoto-1591488320449-011701bb6704?w=1280","2026-07-31T20:04:57.324Z","What is the IQ of Gemini 3.1 Pro?","Gemini 3.1 Pro scores 141 on the Mensa Norway IQ test and 46.5 on the Intelligence Index. We break down every benchmark and explain why it's the most cost-effective frontier model.",{"id":168,"slug":169,"coverImage":170,"author":9,"viewsCount":134,"createdAt":171,"title":172,"description":173,"lang":14,"readingTime":16},"jl69290gswdg1wrb5oy6iiyo","iq-gpt-55","https:\u002F\u002Fimages.unsplash.com\u002Fphoto-1620712943543-bcc4688e7485?w=1280","2026-07-31T19:57:55.350Z","What is the IQ of GPT-5.5?","GPT-5.5 scores ~145 on the Mensa Norway IQ test and 54.8 on the Intelligence Index. We break down every benchmark, explain what the numbers mean, and compare it to Claude, Gemini, and Grok.",{"id":175,"slug":176,"coverImage":177,"author":9,"viewsCount":134,"createdAt":178,"title":179,"description":180,"lang":14,"readingTime":16},"fmyy2tdkvdmffu6n9y5npc5c","iq-claude-opus-48","https:\u002F\u002Fimages.unsplash.com\u002Fphoto-1677442136019-21780ecad995?w=1280","2026-07-31T19:51:43.806Z","What is the IQ of Claude Opus 4.8?","Claude Opus 4.8 tops the Artificial Analysis Intelligence Index at 55.7 and scores ~132-145 on human IQ tests. We break down every benchmark, explain what the numbers mean, and compare it to GPT-5.5, Gemini, and Grok.",{"id":182,"slug":183,"coverImage":184,"author":9,"viewsCount":10,"createdAt":185,"title":186,"description":187,"lang":14,"readingTime":128},"mhlz1hmh5ez4vs6loftlxd06","fasting-cognitive-performance","https:\u002F\u002Fimages.unsplash.com\u002Fphoto-1490645935441-5b6c2cc1e5e1?w=1280","2026-07-30T18:25:44.150Z","Does Fasting Improve Cognitive Performance?","From intermittent fasting to ketones: we review the evidence on whether fasting sharpens your mind, boosts focus, or actually impairs cognition. The science may surprise you.",{"id":189,"slug":190,"coverImage":191,"author":9,"viewsCount":21,"createdAt":192,"title":193,"description":194,"lang":14,"readingTime":16},"fentmgsgc3lq034m2nk53ol4","wechsler-iq-test-subtests","https:\u002F\u002Fimages.unsplash.com\u002Fphoto-1494256997604-668ec142f5d2?w=1280","2026-07-30T18:15:36.675Z","Wechsler IQ Test: What Each Subtest Really Measures","The WAIS and WISC are the most widely used IQ tests in the world. We break down every subtest, what it actually measures, and how scores translate to real cognitive abilities.",{"id":196,"slug":197,"coverImage":198,"author":9,"viewsCount":10,"createdAt":199,"title":200,"description":201,"lang":14,"readingTime":202},"wlwyq33ezei7c22edclwpyvo","lead-fluoride-iq-loss","https:\u002F\u002Fimages.unsplash.com\u002Fphoto-1581093577124-a6f4b6e3a5e1?w=1280","2026-07-30T18:04:10.914Z","Lead, Fluoride, and IQ Loss: The Environmental Evidence","Can environmental toxins actually lower your IQ? We review the strongest studies on lead exposure, fluoride, and cognitive decline — separating solid science from sensationalism.",9,{"id":204,"slug":205,"coverImage":206,"author":9,"viewsCount":31,"createdAt":207,"title":208,"description":209,"lang":14,"readingTime":128},"ue16gxaun3lus9vml8imf6ci","iq-changes-during-pregnancy","https:\u002F\u002Fimages.unsplash.com\u002Fphoto-1559750178-961c031777fe?w=1280","2026-07-29T18:43:46.337Z","IQ Changes During Pregnancy: What Research Finds","Does pregnancy really shrink your brain? We review the neuroscience of maternal cognition, from \"pregnancy brain\" to lasting structural changes, and separate fact from myth.",{"id":211,"slug":212,"coverImage":213,"author":9,"viewsCount":134,"createdAt":214,"title":215,"description":216,"lang":14,"readingTime":202},"zd0sfdazkyai41s0kh639ohl","does-bilingualism-raise-iq","https:\u002F\u002Fimages.unsplash.com\u002Fphoto-1523050854058-8f908275853f?w=1280","2026-07-29T18:29:58.855Z","Does Speaking Two Languages Actually Raise Your IQ?","The bilingual advantage has been debated for decades. We review the strongest studies on bilingualism and cognitive ability, separating real effects from hype.",{"id":218,"slug":219,"coverImage":220,"author":9,"viewsCount":31,"createdAt":221,"title":222,"description":223,"lang":14,"readingTime":202},"s0hko69oudyziplx2q04ab9q","gifted-kids-peaked-early-childhood-iq-fades","https:\u002F\u002Fimages.unsplash.com\u002Fphoto-1503454537195-1dc81782b6d3?w=1280","2026-07-29T18:09:49.555Z","Gifted Kids Who Peaked Early: Why Childhood IQ Fades","Why do so many child prodigies become ordinary adults? We examine the research on early IQ peaking, the \"regression to the mean\" effect, and what actually predicts adult success.",{"id":225,"slug":226,"coverImage":227,"author":9,"viewsCount":228,"createdAt":229,"title":230,"description":231,"lang":14,"readingTime":128},"s7x75o02kixj2u43gw4ah5pa","olfactory-processing-speed-cognitive-marker","https:\u002F\u002Fimages.unsplash.com\u002Fphoto-1542838132-92c53300287c?w=1280",5,"2026-07-28T20:38:15.776Z","Olfactory Processing Speed as a Cognitive Marker","Can how fast you identify smells predict your cognitive ability? We examine the surprising research linking olfactory processing speed to IQ, memory, and neurodegenerative disease risk.",{"id":233,"slug":234,"coverImage":235,"author":9,"viewsCount":21,"createdAt":236,"title":237,"description":238,"lang":14,"readingTime":202},"v1qq1yz082k72zzynk0qv2yh","do-musicians-have-higher-iqs","https:\u002F\u002Fimages.unsplash.com\u002Fphoto-1520006510097-3d64427f3f98?w=1280","2026-07-28T20:25:14.220Z","Do Musicians Have Higher IQs? The Research Verdict","The idea that musicians are smarter is widespread — but does the science support it? We review 10 key studies on musical training, IQ, and cognitive ability to separate fact from fiction.",{"id":240,"slug":241,"coverImage":242,"author":9,"viewsCount":31,"createdAt":243,"title":244,"description":245,"lang":14,"readingTime":202},"xam91ejelkjg23x2z9skt9vm","are-night-owls-smarter","https:\u002F\u002Fimages.unsplash.com\u002Fphoto-1518837695005-2083093ee35b?w=1280","2026-07-28T19:47:32.529Z","Are Night Owls Smarter? Chronotype and Cognitive Performance","The idea that night owls are smarter than early birds is popular — but what does the research actually show? We examine the evidence on chronotype, IQ, and cognitive performance.",{"id":247,"slug":248,"coverImage":249,"author":9,"viewsCount":228,"createdAt":250,"title":251,"description":252,"lang":14,"readingTime":202},"hpakkt3uv7frkkhj2el6fvm1","does-high-iq-make-you-rich","https:\u002F\u002Fimages.unsplash.com\u002Fphoto-1579621970795-87facc2f976d?w=1280","2026-07-28T19:37:46.682Z","Does High IQ Make You Rich? 12 Studies Reviewed","The correlation between IQ and income is real but surprisingly weak. We review 12 key studies to understand why high IQ doesn't guarantee wealth — and what actually predicts financial success.",{"id":254,"slug":255,"coverImage":256,"author":9,"viewsCount":228,"createdAt":257,"title":258,"description":259,"lang":14,"readingTime":202},"b5ct19s9mfso606wcppcp79o","iq-test-scores-in-prison-populations","https:\u002F\u002Fimages.unsplash.com\u002Fphoto-1497366216548-37526070297c?w=1280","2026-07-28T18:14:35.540Z","IQ Test Scores in Prison Populations: What the Studies Show","Research consistently finds that prison populations score lower on IQ tests than the general public — by about 10-12 points on average. But the reasons are far more complex than \"criminals are less intelligent.\" This article examines the evidence, the confounding factors, and what the data actually means.",{"id":261,"slug":262,"coverImage":263,"author":9,"viewsCount":156,"createdAt":264,"title":265,"description":266,"lang":14,"readingTime":16},"zjjd6y8yfcydy6y02c9xa2yf","which-country-has-the-highest-average-iq","https:\u002F\u002Fimages.unsplash.com\u002Fphoto-1526129313847-8d262a6b1f0d?w=1280","2026-07-28T17:59:21.857Z","Which Country Has the Highest Average IQ (And Why the Data Is Messy)","Every few months, a new headline declares that Singapore, Hong Kong, or South Korea has the highest IQ in the world. But behind these rankings lies a tangle of methodological problems, cultural biases, and outdated data. This article unpacks what the research actually shows — and what it doesn't.",{"id":268,"slug":269,"coverImage":270,"author":9,"viewsCount":156,"createdAt":271,"title":272,"description":273,"lang":14,"readingTime":202},"nxfpaqozmm0dpaytyy2cjixv","why-your-iq-test-score-changes-over-time","https:\u002F\u002Fimages.unsplash.com\u002Fphoto-1499750310107-5fef28a66643?w=1280","2026-07-28T17:33:41.377Z","Why Your IQ Test Score Changes Over Time","Your IQ score is not a fixed number stamped on your brain. Research shows that IQ scores can change significantly over a lifetime — sometimes by 20 points or more. This guide explains what drives these changes and what they mean for you.",{"page":134,"pageSize":275,"pageCount":31,"total":276},24,49]