
History of IQ Testing: From the Binet–Simon Scale to Raven's Matrices and Wechsler's Scales
A comprehensive history of intelligence measurement: Binet and mental age, Stern's formula, Stanford–Binet and Army tests, the g-factor, Raven's Progressive Matrices, Wechsler's deviation IQ, myths, and how to read a result.
Why Measure Intelligence at All
The history of IQ testing is not a history of "labeling foreheads." It is a history of attempts to answer a practical question: how to objectively determine who needs additional help in learning, and who needs more challenging tasks.
Before the early 20th century, teachers and physicians relied on subjective impressions: "bright," "slow," "inattentive." Such assessments depended on the adult's mood, the child's language, classroom discipline, and random circumstances. A tool was needed that:
- compares a person against an age-based norm;
- relies on observable tasks rather than "gut feeling";
- serves pedagogy, not just selecting the "best."
It was from this applied task that all modern intelligence psychodiagnostics grew — from the Binet–Simon scale to Raven's matrices and Wechsler's batteries.
1905: Paris, Binet and Simon — the Birth of a Practical Test
In 1905, the French Ministry of Education commissioned psychologist Alfred Binet and physician Théodore Simon to find a way to identify schoolchildren who struggled with the curriculum. The goal was pedagogical: to provide support in time, not to "weed out."
Thus the Binet–Simon scale was born. It included tasks on:
- short-term memory;
- comprehension of instructions;
- attention and discrimination;
- simple reasoning and judgment.
Binet thought not in terms of "permanent IQ" but of mental age: a set of tasks that a typical child of a given age could usually handle. If a child is chronologically 8 years old but consistently performs at the level of a 10-year-old, their mental age is above their chronological one.
1912: William Stern and the IQ Formula
The term IQ (Intelligence Quotient) was proposed by German psychologist William Stern in 1912. He linked mental and chronological age with a simple formula:
IQ = (mental age ÷ chronological age) × 100
Example: a child is 8 years old, mental age by test is 10:
(10 ÷ 8) × 100 = 125
The formula was convenient for schools, but its limits quickly emerged:
- after roughly 16–18 years, "mental age" stops growing as linearly;
- in adults, the "mental age / chronological age" ratio begins to distort meaning;
- a single number masks the profile of strengths and weaknesses.
Nevertheless, it was Stern's idea that cemented the word IQ in both popular culture and science.
1916–1917: Stanford–Binet and Mass Testing
Lewis Terman and Stanford–Binet
In the US, the Binet scale was adapted by Lewis Terman at Stanford (1916). Stanford–Binet became one of the most influential tests of the 20th century: standardization, norms, wide use in schools and research.
It is historically important to remember the shadow of the era: part of early testology intersected with eugenics ideas and rigid social selection. Today this reads as a warning: a measurement instrument can be used ethically or harmfully — responsibility lies with whoever interprets the result.
Army Alpha and Army Beta
In 1917, during World War I, the US conducted mass testing of recruits:
| Test | For Whom | Format |
|---|---|---|
| Army Alpha | literate English speakers | verbal and numerical tasks |
| Army Beta | limited English / low literacy | mostly nonverbal tasks |
The scale was unprecedented: approximately 1.7 million examined. This simultaneously:
- accelerated the development of group testing;
- showed the strengths and weaknesses of "assembly-line" diagnostics;
- intensified debates about cultural fairness of tests.
What They Were Trying to Measure: Spearman's g-Factor
In parallel with practical scales, theory progressed. Charles Spearman proposed the idea of a general factor of intelligence (g): behind different tasks (verbal, numerical, spatial) lies a general capacity for reasoning.
From this came an important conclusion for the history of testing:
- a good IQ test is not a "trivia guessing game";
- it should load on rule extraction, comparison, condition holding, abstraction;
- individual subtests may be verbal or visual, but the goal is to estimate overall cognitive resource.
This is why Raven's matrices were later so highly regarded: they "cleanly" target abstract reasoning with minimal reliance on school knowledge.
1938: John Raven and Progressive Matrices
English psychologist John C. Raven created Raven's Progressive Matrices — a family of tests where one must find the missing piece of a pattern based on the logic of a series or matrix.
Why This Was Revolutionary
- Almost no language. No long verbal instructions within the items themselves.
- Less school knowledge needed. No need to remember dates, formulas, or terms.
- Culturally softer than purely verbal batteries (though no test is truly "culture-free").
- Strong link to fluid intelligence and the g-factor.
Typical variants:
| Version | Usually Suited For | Character |
|---|---|---|
| CPM (Colored) | children, elderly, lighter load | easier entry |
| SPM (Standard) | broad adult range | the classic |
| APM (Advanced) | high ability range | harder ceiling |
1939: David Wechsler and the Modern IQ Scale
David Wechsler changed the very mathematics of interpretation. Instead of ratio-IQ ("mental age / chronological age"), he established the approach:
peer group mean = 100, standard deviation typically = 15
Thus deviation IQ was born: your result is compared against the norm for people your age (and the relevant standardization population), not against an abstract "child's mental age."
The Wechsler Scale Family
| Scale | Audience | Idea |
|---|---|---|
| WAIS | adults | full clinical battery |
| WISC | school-age children | school and clinical diagnostics |
| WPPSI | preschoolers | early profile assessment |
Wechsler batteries do not give a single "magic number" but a profile: verbal comprehension, visual-spatial block, working memory, processing speed, etc. (specific indices depend on the test edition).
Timeline of Intelligence Psychodiagnostics Evolution
| Year / Era | Key Figures | Method | Main Shift |
|---|---|---|---|
| 1905 | Binet, Simon | Binet–Simon scale | Mental age for schools |
| 1912 | Stern | IQ formula | Quotient as ratio of ages |
| 1916 | Terman | Stanford–Binet | Large-scale US standardization |
| 1917 | US military psychologists | Army Alpha / Beta | Mass group testing |
| 1938 | Raven | Progressive Matrices | Nonverbal focus on fluid g |
| 1939+ | Wechsler | WAIS / WISC | Deviation IQ (M=100, SD=15) |
| Late 20th – 21st | Modern editions | Updated norms, computerization | Adaptive testing, online formats |
Ratio IQ vs Deviation IQ: What's the Difference in Plain Terms
| Ratio IQ (Stern) | Deviation IQ (Wechsler onward) | |
|---|---|---|
| Meaning | ratio of "mental" and chronological age | position relative to peer norm |
| Mean | roughly around 100 when ages match | fixed at 100 by standardization |
| Adults | formula breaks down | works through age-based norms |
| Interpretation | "how many years ahead/behind" | "how many SD from the mean" |
Example deviations:
- ≈ 100 — near average;
- ≈ 115 — roughly +1 SD (above average);
- ≈ 85 — roughly −1 SD (below average);
- extreme tails are rare and require especially careful interpretation.
<WARNING_BLOCK title="Don't Confuse "Online Score" with Clinical Diagnosis">A short web practice is useful for training and orientation. A proper clinical/educational assessment is a standardized procedure with norms, professional qualifications, and ethics. NeuroLab is a training and educational environment, not a replacement for official diagnostics.
What Good IQ Tests Usually Measure (and What They Don't)
Typically Load On
- Fluid reasoning — reasoning on novel material;
- Crystallized knowledge (in verbal batteries) — vocabulary, information;
- Working memory — holding and updating information;
- Processing speed — rate of simple cognitive operations;
- Visual-spatial analysis — patterns, assembly, orientation.
Usually Don't Measure Directly
- moral qualities and character;
- creativity in all its forms;
- motivation, sleep, anxiety, and health at test time;
- narrow specialized professional skill;
- "life success" as a single quantity.
Cultural Fairness: The Myth of a "Completely Neutral" Test
Raven's matrices are often called "culture-free." It is more accurate to say: less dependent on language and school knowledge. But factors remain:
- familiarity with the test format ("test-wiseness");
- quality of instructions and a calm environment;
- vision, fatigue, time anxiety;
- experience with abstract patterns (education, games, STEM).
So an honest history of IQ includes critique as well: tests improved fairness relative to pure teacher subjectivity, but were never a magic mirror of "pure brain outside culture."
How to Read a Result: A Short Common-Sense Guide
- Look at the range and repetition, not a single attempt.
- Factor in your state: sleep deprivation easily "eats" tens of minutes of concentration.
- If there is a subtest profile — look for strengths, not just the "hole."
- Use the result as a learning navigator: where to train reasoning, memory, speed, attention.
- Don't turn the score into an identity ("I = a number").
Mini Checklist Before a Serious Attempt
- slept as well as possible;
- notifications off;
- understood the instructions before starting;
- not taking the test "angrily" after a hard day;
- ready to accept the result as a starting point, not a verdict.
Connection to Other Cognitive Domains
The history of IQ testing logically aligns with what NeuroLab trains alongside:
| Domain | Connection to IQ History | Why Practice Alongside |
|---|---|---|
| Matrix practice | the Raven / fluid g line | rule extraction |
| Reaction | processing speed | less "fog" at the start |
| Schulte / attention | search control | resilience to visual noise |
| Working memory | part of Wechsler batteries | holding task conditions |
Common Myths About the History and Meaning of IQ
"Binet wanted to rank people permanently."
No: the original commission was about pedagogical help for children with learning difficulties.
"IQ = genetics, period."
Heritability of cognitive measures is researched, but environment, education, health, and opportunity strongly influence the realization of potential. It is not "either/or."
"Raven completely eliminates culture."
Reduces language load — yes. Eliminates culture — no.
"Wechsler just renamed Binet."
No: the key shift is deviation IQ and the expanded profile approach.
"An online test equals a clinical WAIS."
Formats may be related in task concept, but norms, length, proctoring, and purpose differ.
What Changed in the 21st Century
Modern intelligence assessment increasingly includes:
- computerized and adaptive testing (difficulty adjusts to responses);
- updating norms for new generations;
- interest in ability profiles, not just a single score;
- combining classic tasks with metrics of attention, speed, and stability.
Online platforms like NeuroLab stand in this line as an accessible practice ground: you can regularly train reasoning and related skills, see dynamics, and not wait for "one fateful commission."
Brief Summary
The history of IQ is a journey from the Paris schoolroom Binet–Simon scale to Stern's formula, mass military testing, Raven's nonverbal matrices, and Wechsler's deviation scales.
Remember the framework:
- Binet — help with learning through mental age;
- Stern — the IQ quotient as a ratio of ages;
- Terman / Army — scaling and standardization (with ethical lessons);
- Raven — visual reasoning with less language load;
- Wechsler — comparison against peer norms and profiling;
- Today — practice, repeated measurement, and sober interpretation matter more than the cult of a single number.
If you want not only to know the history but to feel the logic of the tasks yourself — start with the NeuroLab matrix module and look at the trend across several attempts, not one random run.


