What is the IQ of Qwen 3.5?
Dr. Daria Dudnik 10 min read 1 views

What is the IQ of Qwen 3.5?

Qwen 3.5 scores 115 on the Mensa Norway IQ test and ranks in the top 10 on the Intelligence Index. We break down every benchmark and explain what makes Alibaba's model unique.

Qwen 3.5: Alibaba's versatile contender

Qwen 3.5 is the flagship model from Alibaba's DAMO Academy, one of China's largest and most well-funded AI research labs. Released in early 2026, Qwen 3.5 scores 115 on the Mensa Norway IQ test and ranks in the top 10 on the Artificial Analysis Intelligence Index v4.1. This places it above DeepSeek V4 Pro (111) and competitive with Kimi K3 (118), but below the frontier models (Claude ~130, Gemini 141, GPT-5.5 ~145, Grok-4.20 145).

What makes Qwen 3.5 unique is its versatility. Unlike DeepSeek (which focuses on cost), Kimi (which focuses on context length), or Grok (which focuses on raw IQ), Qwen 3.5 aims to be a well-rounded model that excels across a wide range of tasks. It has strong multilingual support (27+ languages), excellent multimodal capabilities (text, image, audio, video), and competitive pricing — making it the Swiss Army knife of AI models.

Qwen 3.5 is the all-rounder: a model that may not be the absolute best at any single thing, but is very good at almost everything. For organizations that need a single model for diverse workloads — customer service, content generation, code analysis, translation, visual understanding — Qwen 3.5 offers an compelling combination of capability, versatility, and value.

Key scores:

  • Artificial Analysis Intelligence Index: top 10 (score ~45-46)
  • Mensa Norway IQ (TrackingAI): 115 (above average)
  • AI IQ composite (VexoWire): ~115 estimated
  • GDPval-AA v2: top 10
  • Humanity's Last Exam: top 10
  • Hallucination rate: low-moderate
  • Cost: 0.50/1.50 per 1M tokens (very affordable)
  • Multimodal: text, image, audio, video
  • Languages: 27+ languages supported

1. The versatility advantage

Qwen 3.5 represents Alibaba's bet that the future of AI is not about being the absolute smartest, but about being the most versatile. While other labs optimize for specific dimensions — IQ, cost, context, personality — Alibaba has built a model that is competitive across all of them.

Consider the profile:

  • IQ: 115 (above average, competitive with other Chinese models)
  • Cost: 0.50/1.50 per 1M tokens (among the cheapest)
  • Multimodal: text, image, audio, video (full multimodal)
  • Languages: 27+ languages (best-in-class multilingual)
  • Context: 256K tokens (respectable, though not Kimi-level)
  • Open-source: weights available for local deployment

No single attribute is the best in class, but the combination is unmatched. DeepSeek is cheaper but lacks multimodal. Kimi has more context but fewer languages. Gemini is smarter but more expensive. Claude is more capable but closed-source. Qwen 3.5 threads the needle — good enough at everything to be the one model you need.

This versatility makes Qwen 3.5 particularly attractive for:

  • Enterprise deployments with diverse workloads
  • International applications requiring many languages
  • Multimodal applications combining text, image, and audio
  • Cost-conscious organizations that still need quality
  • Self-hosted deployments via open-source weights

2. Benchmark breakdown

Mensa Norway IQ test: 115 (above average)

The Mensa Norway IQ test consists of 36 visual pattern recognition questions. Qwen 3.5 scores 115 — above the human average of 100, placing it in the "above average" range. This is higher than DeepSeek V4 Pro (111), slightly below Kimi K3 (118), and below the frontier models (Claude ~130, Gemini 141, GPT-5.5 ~145, Grok-4.20 145).

An IQ of 115 means Qwen 3.5 is smarter than approximately 84% of humans — the top 16% of human intelligence. If it were a person, it would be comparable to a strong university student or a capable professional. Not a genius, but certainly bright and competent.

Artificial Analysis Intelligence Index: top 10

On the Intelligence Index, Qwen 3.5 ranks in the top 10 globally with an estimated score of ~45-46. This places it competitive with Gemini 3.1 Pro (46.5), Kimi K3 (~45-47), and DeepSeek V4 Pro (44.3), but behind Claude (55.7) and GPT-5.5 (54.8).

Breaking down the Index components:

  • GDPval-AA v2: Qwen 3.5 performs well, benefiting from its multimodal capabilities for diverse agentic tasks
  • AA-Omniscience: Qwen 3.5 has a low-moderate hallucination rate, benefiting from Alibaba's extensive training data
  • GPQA: Qwen 3.5 performs adequately on graduate-level science questions
  • HLE: Qwen 3.5 ranks in the top 10 on Humanity's Last Exam
  • ARC-AGI: Qwen 3.5 performs reasonably on abstract reasoning
  • LMSYS Arena: Qwen 3.5 gets strong human preference scores, particularly for multilingual and multimodal tasks

AI IQ composite (VexoWire): ~115

The VexoWire composite IQ blends multiple tests (Mensa Norway, Mensa Denmark, Raven's, WAIS-IV). Qwen 3.5's estimated composite IQ is ~115 — consistent with its Mensa Norway score. This places it at the boundary between "average" and "above average," smarter than about 84% of humans.

GDPval-AA v2: top 10

On the agentic benchmark, Qwen 3.5 ranks in the top 10. Its multimodal capabilities give it an advantage on agentic tasks that involve processing different types of data — text, images, and audio. For moderate agentic workloads, Qwen 3.5 is competent, though it trails Claude and GPT-5.5 on the most complex tasks.

Humanity's Last Exam (HLE): top 10

On the hardest academic questions, Qwen 3.5 ranks in the top 10 among AI models. Alibaba's strong academic training data and the model's broad knowledge base help it perform well across disciplines.


3. What does an IQ of 115 mean?

To put Qwen 3.5's IQ of 115 in context:

  • 85-115 (68% of people): Average intelligence
  • 115-130 (14%): Above average to high average
  • 130-145 (2%): Gifted — qualifies for Mensa (top 2%)
  • 145-160 (0.1%): Highly gifted / genius

Qwen 3.5, at 115, sits right at the boundary between "average" and "above average" — smarter than about 84% of humans. If it were a person, it would be comparable to a strong university student or a capable professional. Not a genius, but certainly bright and competent.

This is higher than DeepSeek V4 Pro (111) but below the frontier models:

  • Claude Opus 4.8: ~130 (gifted)
  • Gemini 3.1 Pro: 141 (highly gifted)
  • GPT-5.5: ~145 (genius)
  • Grok-4.20: 145 (genius)

The gap to the frontier models is about 15-30 IQ points. But Qwen 3.5 compensates with its versatility — for tasks that require multimodal understanding, multilingual support, or diverse workloads, Qwen 3.5's all-around capabilities can outweigh the IQ gap.


4. Where Qwen 3.5 excels

Multilingual support

Qwen 3.5 supports 27+ languages natively, making it one of the most multilingual AI models available. This isn't just translation — the model is genuinely trained in these languages, understanding cultural context, idioms, and nuances. For international applications, Qwen 3.5 is among the best choices.

Multimodal capabilities

Qwen 3.5 can process text, images, audio, and video — making it a true multimodal model. This enables applications that other text-only models cannot handle: visual question answering, image analysis, audio transcription, video understanding, and more. For applications that combine multiple media types, Qwen 3.5 is exceptionally capable.

Cost-effectiveness

At 0.50/1.50 per 1M tokens, Qwen 3.5 is very affordable — slightly more expensive than DeepSeek V4 Pro (0.44/0.87) but significantly cheaper than Claude (5/25), GPT-5.5 (5/30), and Gemini (2/12). Given its multimodal capabilities and multilingual support, the value proposition is excellent.

Open-source availability

Qwen 3.5's weights are available for local deployment, making it one of the few multimodal open-source models. This is significant — most multimodal models (GPT-5.5, Gemini) are closed-source. For organizations that need multimodal capabilities with data privacy, Qwen 3.5 is one of the best options.

Chinese-language tasks

As a Chinese-developed model, Qwen 3.5 has excellent Chinese-language capabilities. For Chinese-language applications, Qwen 3.5 is among the best models available, often outperforming Western models on Chinese-specific tasks.

Enterprise readiness

Alibaba's infrastructure and enterprise experience make Qwen 3.5 particularly well-suited for enterprise deployments. The model is available through Alibaba Cloud with enterprise-grade SLAs, compliance certifications, and integration support. For businesses looking for a reliable, versatile AI model, Qwen 3.5 is a strong choice.


5. Where Qwen 3.5 falls short

Raw IQ and complex reasoning

With an IQ of 115, Qwen 3.5 is below the frontier models on complex reasoning tasks. For the hardest problems — competition mathematics, advanced logic puzzles, cutting-edge scientific reasoning — Claude, GPT-5.5, and Gemini are substantially better. If your task requires genius-level reasoning, the frontier models are the better choice.

Context window

With a 256K token context window, Qwen 3.5 is adequate but not exceptional. Kimi K3 (2M) and Gemini (1M) offer significantly larger context windows. For tasks that require processing very long documents in a single pass, Kimi K3 or Gemini are better choices.

Agentic task complexity

While Qwen 3.5's multimodal capabilities help with diverse agentic tasks, it trails Claude and GPT-5.5 on pure task complexity. For the most demanding agentic applications — writing complex software from scratch, managing intricate research workflows — Claude remains more capable.

Emotional intelligence

Qwen 3.5's EQ (emotional intelligence) is estimated at ~107 — lower than Claude (~132), GPT-5.5 (~120), and Gemini (~115), but comparable to DeepSeek (~105) and Grok (~110). Its responses tend to be more functional and less emotionally nuanced. For therapy apps, customer service, or sensitive human interactions, Claude is the better choice.

Hallucination rate

While Qwen 3.5 has a low-moderate hallucination rate, it's not as low as Claude (35.9%). For applications where factual accuracy is critical — research, law, medicine — Claude remains more reliable.


6. Qwen 3.5 vs other models

Model Intelligence Index Mensa Norway IQ AI IQ est. Multimodal Cost/1M Languages Open-source
Qwen 3.5 ~45-46 (top 10) 115 ~115 text+image+audio+video 0.50/1.50 27+ yes
Claude Opus 4.8 55.7 ~130 ~132 text+image 5/25 ~20 no
GPT-5.5 54.8 ~145 ~136 text+image+audio 5/30 ~50 no
Gemini 3.1 Pro 46.5 141 ~131 text+image+audio+video 2/12 ~40 no
Grok-4.20 top 5 145 ~140 text+image 3/15 ~20 no
DeepSeek V4 Pro 44.3 ~111 ~111 text only 0.44/0.87 ~10 yes
Kimi K3 ~45-47 (top 10) 118 ~118 text only 0.55/2.19 ~10 no

The picture: Qwen 3.5 is the all-rounder — not the best at any single thing, but very good at almost everything. Claude is the capable professional — best at real-world tasks and factual accuracy. GPT-5.5 is the test genius — best at math and visual reasoning. Gemini is the value champion — best cost-effectiveness among frontier models. Grok-4.20 is the raw IQ champion — highest IQ with personality. DeepSeek is the budget champion — cheapest with open-source. And Kimi is the long-context champion — biggest context window.

For organizations that need versatility — multimodal, multilingual, affordable, open-source — Qwen 3.5 is the clear choice. It's the Swiss Army knife of AI models.


7. What does this mean for humans?

Qwen 3.5 operates at an IQ of ~115 — smarter than about 84% of humans. But:

  • IQ isn't everything: an IQ of 115 is above average and sufficient for most practical tasks
  • Versatility matters more than raw IQ for many applications: for diverse workloads, Qwen 3.5's all-around capabilities are more valuable than a higher IQ in a narrower model
  • Multimodal is the future: Qwen 3.5's ability to process text, images, audio, and video opens up applications that text-only models cannot handle
  • Multilingual enables global reach: with 27+ languages, Qwen 3.5 can serve users in their native language — a capability that matters for global applications
  • Knowledge ≠ understanding: like all AI models, Qwen 3.5 has access to vast knowledge but doesn't truly "understand" the way a human does
  • No consciousness: high IQ doesn't mean sentience. Qwen 3.5 is a sophisticated pattern-matching system, not a thinking being
  • Open-source empowers innovation: Qwen 3.5's open weights allow researchers and developers to build custom solutions, fine-tune for specific domains, and run locally

The most honest answer: Qwen 3.5 is smarter than 84% of humans on cognitive tests, and for applications that require versatility — multimodal, multilingual, affordable — it is the smartest choice — not because it's the most intelligent, but because it's the most capable across the broadest range of tasks.

How does your IQ compare to Qwen 3.5? Take NeuroLab's professional IQ assessment to find out.
Take the Test Now →


Conclusion

💡 Key takeaways
- Qwen 3.5's IQ is 115 on the Mensa Norway test — above average, smarter than ~84% of humans.

  • Top 10 on the Intelligence Index — competitive with Gemini, Kimi, and DeepSeek, behind Claude and GPT-5.5.
  • Full multimodal — text, image, audio, and video processing in a single model.
  • 27+ languages — one of the most multilingual AI models available.
  • Open-source — weights available for local deployment, enabling privacy and customization.
  • Best for: diverse enterprise workloads, international applications, multimodal tasks, cost-conscious organizations, self-hosted deployments.
  • Not best for: raw IQ (frontier models win), long context (Kimi wins), complex agentic tasks (Claude wins), emotional intelligence (Claude wins).
  • The lesson: intelligence isn't just about being the smartest — it's about being the most versatile. Qwen 3.5 proves that being good at everything can be more valuable than being the best at one thing.