What is the IQ of Kimi K3?
Dr. Daria Dudnik 10 min read 2 views

What is the IQ of Kimi K3?

Kimi K3 scores 118 on the Mensa Norway IQ test and ranks in the top 10 on the Intelligence Index. We break down every benchmark and explain what makes Moonshot AI's model unique.

Kimi K3: Moonshot AI's rising star

Kimi K3 is the flagship model from Moonshot AI, one of China's most innovative AI startups. Released in early 2026, Kimi K3 scores 118 on the Mensa Norway IQ test and ranks in the top 10 on the Artificial Analysis Intelligence Index v4.1. This places it above DeepSeek V4 Pro (111) but below the frontier models (Claude ~130, Gemini 141, GPT-5.5 ~145, Grok-4.20 145).

What makes Kimi K3 unique is its extraordinary context window. While most AI models process 32K-200K tokens at a time, Kimi K3 can handle up to 2 million tokens in a single context — roughly 1,500 pages of text, or about 5-10 full books. This makes it the undisputed champion of long-context tasks: analyzing entire codebases, processing lengthy legal documents, summarizing book series, and understanding complex multi-document relationships.

Kimi K3 is the long-context specialist: a model that may not match the frontier on raw IQ, but can process and reason over amounts of information that other models simply cannot handle. For applications that require understanding vast amounts of text in a single pass, Kimi K3 is in a league of its own.

Key scores:

  • Artificial Analysis Intelligence Index: top 10 (score ~45-47)
  • Mensa Norway IQ (TrackingAI): 118 (above average, approaching gifted)
  • AI IQ composite (VexoWire): ~118 estimated
  • GDPval-AA v2: top 10
  • Humanity's Last Exam: top 10
  • Hallucination rate: low-moderate
  • Cost: 0.55/2.19 per 1M tokens (very affordable)
  • Context window: 2 million tokens (largest of any major model)

1. The long-context revolution

Kimi K3 represents a different approach to AI capability. While most labs focus on making models smarter (higher IQ, better reasoning), Moonshot AI has focused on making models able to process more information at once. The result is a model with a 2-million-token context window — 10-60x larger than most competitors.

To put this in perspective:

  • GPT-5.5: ~200K token context
  • Claude Opus 4.8: ~200K token context
  • Gemini 3.1 Pro: ~1M token context
  • Kimi K3: 2M token context

With 2M tokens, Kimi K3 can process:

  • An entire software codebase (50,000+ lines of code)
  • 5-10 full-length novels in a single prompt
  • A complete legal case file with all evidence and precedents
  • An entire research paper collection on a topic
  • Hours of transcribed meeting recordings

This isn't just about reading more — it's about understanding relationships across documents. Kimi K3 can identify connections between page 1 and page 1,500 of a document that other models would lose track of. For tasks that require holistic understanding of large information sets, this is a game-changer.


2. Benchmark breakdown

Mensa Norway IQ test: 118 (above average)

The Mensa Norway IQ test consists of 36 visual pattern recognition questions. Kimi K3 scores 118 — above the human average of 100 and approaching the "gifted" threshold of 130. This is higher than DeepSeek V4 Pro (111) but below the frontier models (Claude ~130, Gemini 141, GPT-5.5 ~145, Grok-4.20 145).

An IQ of 118 means Kimi K3 is smarter than approximately 85% of humans. If it were a person, it would be comparable to a strong university graduate — bright and capable, though not at the genius level of the frontier models. For most practical reasoning tasks, this level of intelligence is more than sufficient.

Artificial Analysis Intelligence Index: top 10

On the Intelligence Index, Kimi K3 ranks in the top 10 globally with an estimated score of ~45-47. This places it competitive with Gemini 3.1 Pro (46.5) and DeepSeek V4 Pro (44.3), but behind Claude (55.7) and GPT-5.5 (54.8).

Breaking down the Index components:

  • GDPval-AA v2: Kimi K3 performs well, benefiting from its long context for multi-step tasks
  • AA-Omniscience: Kimi K3 has a low-moderate hallucination rate, benefiting from its ability to cross-reference information within its context
  • GPQA: Kimi K3 performs adequately on graduate-level science questions
  • HLE: Kimi K3 ranks in the top 10 on Humanity's Last Exam
  • ARC-AGI: Kimi K3 performs reasonably on abstract reasoning
  • LMSYS Arena: Kimi K3 gets strong human preference scores, particularly for long-document tasks

AI IQ composite (VexoWire): ~118

The VexoWire composite IQ blends multiple tests (Mensa Norway, Mensa Denmark, Raven's, WAIS-IV). Kimi K3's estimated composite IQ is ~118 — consistent with its Mensa Norway score. This places it in the "above average" range, approaching but not reaching the "gifted" level.

GDPval-AA v2: top 10

On the agentic benchmark, Kimi K3 ranks in the top 10. Its long context window gives it a unique advantage on agentic tasks that involve processing large amounts of information — it can hold entire project contexts in memory, making it better at multi-step tasks that require sustained context. However, on pure reasoning complexity, it still trails Claude and GPT-5.5.

Humanity's Last Exam (HLE): top 10

On the hardest academic questions, Kimi K3 ranks in the top 10 among AI models. Its ability to process large amounts of reference material in context gives it an advantage on questions that require synthesizing information from multiple sources.


3. What does an IQ of 118 mean?

To put Kimi K3's IQ of 118 in context:

  • 85-115 (68% of people): Average intelligence
  • 115-130 (14%): Above average to high average
  • 130-145 (2%): Gifted — qualifies for Mensa (top 2%)
  • 145-160 (0.1%): Highly gifted / genius

Kimi K3, at 118, sits in the "above average" range — smarter than about 85% of humans. If it were a person, it would be comparable to a strong university graduate or a capable professional. Not a genius, but certainly bright and competent.

This is higher than DeepSeek V4 Pro (111) but below the frontier models:

  • Claude Opus 4.8: ~130 (gifted)
  • Gemini 3.1 Pro: 141 (highly gifted)
  • GPT-5.5: ~145 (genius)
  • Grok-4.20: 145 (genius)

The gap to the frontier models is about 12-27 IQ points. But Kimi K3 compensates for its lower raw IQ with its extraordinary context window — for tasks that require processing large amounts of information, Kimi K3's long-context advantage can outweigh the IQ gap.


4. Where Kimi K3 excels

Long-context processing

This is Kimi K3's defining feature. With a 2-million-token context window, it can process and reason over amounts of text that no other major model can handle. For analyzing entire codebases, processing lengthy legal documents, summarizing book collections, or understanding complex multi-document relationships, Kimi K3 is the best model available.

Code analysis

Kimi K3 is particularly strong at code analysis tasks. Its long context allows it to understand entire software projects — not just individual files — making it excellent for code review, refactoring suggestions, and architectural analysis. For developers working with large codebases, Kimi K3 offers capabilities that other models simply cannot match.

Document summarization

For summarizing long documents — legal cases, research papers, technical specifications, financial reports — Kimi K3 excels. Its ability to process the entire document at once means its summaries capture nuances and connections that chunked processing might miss.

Cost-effectiveness

At 0.55/2.19 per 1M tokens, Kimi K3 is very affordable — more expensive than DeepSeek V4 Pro (0.44/0.87) but significantly cheaper than Claude (5/25), GPT-5.5 (5/30), and Gemini (2/12). For long-context tasks, the value proposition is exceptional: you get 2M token context at a fraction of the cost of frontier models.

Chinese-language tasks

As a Chinese-developed model, Kimi K3 has excellent Chinese-language capabilities. For Chinese-language applications — content generation, analysis, translation — Kimi K3 is among the best models available, sometimes outperforming Western models.

Research assistance

For researchers who need to process large amounts of literature, Kimi K3 is an invaluable tool. It can ingest dozens of papers at once and identify connections, contradictions, and gaps across them — a task that would take a human researcher days or weeks.


5. Where Kimi K3 falls short

Raw IQ and complex reasoning

With an IQ of 118, Kimi K3 is below the frontier models on complex reasoning tasks. For the hardest problems — competition mathematics, advanced logic puzzles, cutting-edge scientific reasoning — Claude, GPT-5.5, and Gemini are substantially better. If your task requires genius-level reasoning on short inputs, the frontier models are the better choice.

Agentic task complexity

While Kimi K3's long context helps with agentic tasks, it still trails Claude and GPT-5.5 on pure task complexity. For the most demanding agentic applications — writing complex software from scratch, managing intricate research workflows — Claude remains more capable.

Emotional intelligence

Kimi K3's EQ (emotional intelligence) is estimated at ~108 — lower than Claude (~132), GPT-5.5 (~120), and Gemini (~115), but comparable to Grok (~110). Its responses tend to be more functional and less emotionally nuanced. For therapy apps, customer service, or sensitive human interactions, Claude is the better choice.

Global availability

Kimi K3 is primarily available through Moonshot AI's API and select partners. It has less global distribution than Claude, GPT-5.5, or Gemini, which may affect availability in certain regions. However, it is increasingly available through third-party platforms.

Brand recognition

As a relatively new model from a Chinese startup, Kimi K3 has less brand recognition and trust than models from OpenAI, Anthropic, or Google. For organizations that prioritize working with established Western providers, this may be a consideration.


6. Kimi K3 vs other models

Model Intelligence Index Mensa Norway IQ AI IQ est. Context Cost/1M Best for
Kimi K3 ~45-47 (top 10) 118 ~118 2M 0.55/2.19 Long-context tasks
Claude Opus 4.8 55.7 ~130 ~132 200K 5/25 Real-world tasks
GPT-5.5 54.8 ~145 ~136 200K 5/30 Math & reasoning
Gemini 3.1 Pro 46.5 141 ~131 1M 2/12 Value & multimodal
Grok-4.20 top 5 145 ~140 200K 3/15 Raw IQ & personality
DeepSeek V4 Pro 44.3 ~111 ~111 128K 0.44/0.87 Budget & open-source

The picture: Kimi K3 is the long-context champion — not the smartest, but able to process more information at once than any other model. Claude is the capable professional — best at real-world tasks and factual accuracy. GPT-5.5 is the test genius — best at math and visual reasoning. Gemini is the value champion — best cost-effectiveness among frontier models. Grok-4.20 is the raw IQ champion — highest IQ with personality. And DeepSeek is the budget champion — cheapest with open-source weights.

For users who need to process vast amounts of text in a single pass, Kimi K3 is the clear choice — its 2M context window is in a league of its own.


7. What does this mean for humans?

Kimi K3 operates at an IQ of ~118 — smarter than about 85% of humans. But:

  • IQ isn't everything: an IQ of 118 is above average and sufficient for most practical tasks
  • Context matters more than IQ for many tasks: for tasks involving large documents, Kimi K3's 2M context window is more valuable than a higher IQ
  • Long-context enables new applications: Kimi K3 opens up possibilities that other models cannot — analyzing entire codebases, processing book-length documents, synthesizing research across dozens of papers
  • Knowledge ≠ understanding: like all AI models, Kimi K3 has access to vast knowledge but doesn't truly "understand" the way a human does
  • No consciousness: high IQ doesn't mean sentience. Kimi K3 is a sophisticated pattern-matching system, not a thinking being
  • Specialization vs generalization: Kimi K3 proves that AI can compete by being the best at a specific capability (long context) rather than trying to be the best at everything

The most honest answer: Kimi K3 is smarter than 85% of humans on cognitive tests, and for tasks that require processing large amounts of information, it is the smartest choice — not because it's the most intelligent, but because it can hold more in its "mind" at once.

How does your IQ compare to Kimi K3? Take NeuroLab's professional IQ assessment to find out.
Take the Test Now →


Conclusion

💡 Key takeaways
- Kimi K3's IQ is 118 on the Mensa Norway test — above average, smarter than ~85% of humans.

  • Top 10 on the Intelligence Index — competitive with Gemini and DeepSeek, behind Claude and GPT-5.5.
  • 2-million-token context window — the largest of any major model, 10-60x larger than most competitors.
  • Best for: long-context tasks, code analysis, document summarization, research assistance, Chinese-language applications.
  • Not best for: raw IQ (frontier models win), complex agentic tasks (Claude wins), emotional intelligence (Claude wins).
  • The lesson: intelligence isn't just about being the smartest — it's about being able to process the most information. Kimi K3 proves that context is a form of intelligence.