History of IQ Testing: From the Binet–Simon Scale to Raven's Matrices and Wechsler's Scales
Dr. Daria Dudnik 13 min read 5 views

History of IQ Testing: From the Binet–Simon Scale to Raven's Matrices and Wechsler's Scales

A comprehensive history of intelligence measurement: Binet and mental age, Stern's formula, Stanford–Binet and Army tests, the g-factor, Raven's Progressive Matrices, Wechsler's deviation IQ, myths, and how to read a result.

Why Measure Intelligence at All

The history of IQ testing is not a history of "labeling foreheads." It is a history of attempts to answer a practical question: how to objectively determine who needs additional help in learning, and who needs more challenging tasks.

Before the early 20th century, teachers and physicians relied on subjective impressions: "bright," "slow," "inattentive." Such assessments depended on the adult's mood, the child's language, classroom discipline, and random circumstances. A tool was needed that:

  1. compares a person against an age-based norm;
  2. relies on observable tasks rather than "gut feeling";
  3. serves pedagogy, not just selecting the "best."

It was from this applied task that all modern intelligence psychodiagnostics grew — from the Binet–Simon scale to Raven's matrices and Wechsler's batteries.

💡 An Important Reading Frame
IQ is a measure of performance on standardized tasks at the time of testing — not a verdict on personality, not a measure of "human worth," and not a complete portrait of talent. A good test is useful if interpreted correctly.


1905: Paris, Binet and Simon — the Birth of a Practical Test

In 1905, the French Ministry of Education commissioned psychologist Alfred Binet and physician Théodore Simon to find a way to identify schoolchildren who struggled with the curriculum. The goal was pedagogical: to provide support in time, not to "weed out."

Thus the Binet–Simon scale was born. It included tasks on:

  • short-term memory;
  • comprehension of instructions;
  • attention and discrimination;
  • simple reasoning and judgment.

Binet thought not in terms of "permanent IQ" but of mental age: a set of tasks that a typical child of a given age could usually handle. If a child is chronologically 8 years old but consistently performs at the level of a 10-year-old, their mental age is above their chronological one.

💡 Alfred Binet's Core Idea
Binet explicitly warned against fatalism. Intelligence was not, for him, a fixed quantity "from nature forever." The test was meant to intervene in time: change the teaching, the load, the support — not slap on a label.

1905
The year of the first practical Binet–Simon scale — the starting point of modern intelligence testology.


1912: William Stern and the IQ Formula

The term IQ (Intelligence Quotient) was proposed by German psychologist William Stern in 1912. He linked mental and chronological age with a simple formula:

IQ = (mental age ÷ chronological age) × 100

Example: a child is 8 years old, mental age by test is 10:

(10 ÷ 8) × 100 = 125

The formula was convenient for schools, but its limits quickly emerged:

  • after roughly 16–18 years, "mental age" stops growing as linearly;
  • in adults, the "mental age / chronological age" ratio begins to distort meaning;
  • a single number masks the profile of strengths and weaknesses.

Nevertheless, it was Stern's idea that cemented the word IQ in both popular culture and science.

⚠️ The Limitation of Ratio IQ
The classic "age-based" formula works well in early-20th-century childhood education and poorly as a universal metric for adults. Modern scales have almost universally moved to deviation IQ (deviation from the peer norm).


1916–1917: Stanford–Binet and Mass Testing

Lewis Terman and Stanford–Binet

In the US, the Binet scale was adapted by Lewis Terman at Stanford (1916). Stanford–Binet became one of the most influential tests of the 20th century: standardization, norms, wide use in schools and research.

It is historically important to remember the shadow of the era: part of early testology intersected with eugenics ideas and rigid social selection. Today this reads as a warning: a measurement instrument can be used ethically or harmfully — responsibility lies with whoever interprets the result.

Army Alpha and Army Beta

In 1917, during World War I, the US conducted mass testing of recruits:

Test For Whom Format
Army Alpha literate English speakers verbal and numerical tasks
Army Beta limited English / low literacy mostly nonverbal tasks

The scale was unprecedented: approximately 1.7 million examined. This simultaneously:

  • accelerated the development of group testing;
  • showed the strengths and weaknesses of "assembly-line" diagnostics;
  • intensified debates about cultural fairness of tests.

1.7 million
Approximately this many US recruits took Army Alpha/Beta — the first giant experiment in mass psychometrics.


What They Were Trying to Measure: Spearman's g-Factor

In parallel with practical scales, theory progressed. Charles Spearman proposed the idea of a general factor of intelligence (g): behind different tasks (verbal, numerical, spatial) lies a general capacity for reasoning.

From this came an important conclusion for the history of testing:

  • a good IQ test is not a "trivia guessing game";
  • it should load on rule extraction, comparison, condition holding, abstraction;
  • individual subtests may be verbal or visual, but the goal is to estimate overall cognitive resource.

This is why Raven's matrices were later so highly regarded: they "cleanly" target abstract reasoning with minimal reliance on school knowledge.


1938: John Raven and Progressive Matrices

English psychologist John C. Raven created Raven's Progressive Matrices — a family of tests where one must find the missing piece of a pattern based on the logic of a series or matrix.

Why This Was Revolutionary

  1. Almost no language. No long verbal instructions within the items themselves.
  2. Less school knowledge needed. No need to remember dates, formulas, or terms.
  3. Culturally softer than purely verbal batteries (though no test is truly "culture-free").
  4. Strong link to fluid intelligence and the g-factor.

Typical variants:

Version Usually Suited For Character
CPM (Colored) children, elderly, lighter load easier entry
SPM (Standard) broad adult range the classic
APM (Advanced) high ability range harder ceiling

💡 What the Matrix Format Trains
You learn to quickly test hypotheses: "does the number of elements change? does the shape rotate? does the fill alternate?" This is a skill of guided rule search, not memorizing answers.

~0.7–0.8+
The typical range of correlations between good matrix tests and g / fluid reasoning estimates in research (depends on sample and battery). This is one reason for the "cult of Raven" in psychometrics.

Take a nonverbal-matrix-style practice session on NeuroLab: no unnecessary "trivia," with a focus on rule extraction and result feedback.
Take the Test Now →


1939: David Wechsler and the Modern IQ Scale

David Wechsler changed the very mathematics of interpretation. Instead of ratio-IQ ("mental age / chronological age"), he established the approach:

peer group mean = 100, standard deviation typically = 15

Thus deviation IQ was born: your result is compared against the norm for people your age (and the relevant standardization population), not against an abstract "child's mental age."

The Wechsler Scale Family

Scale Audience Idea
WAIS adults full clinical battery
WISC school-age children school and clinical diagnostics
WPPSI preschoolers early profile assessment

Wechsler batteries do not give a single "magic number" but a profile: verbal comprehension, visual-spatial block, working memory, processing speed, etc. (specific indices depend on the test edition).

💡 Why the Profile Matters More Than a Single Number
Two people with IQ ≈ 110 can have completely different "pictures" of abilities: one stronger in speech and knowledge, the other in visual reasoning and speed. For study and career, the profile is often more useful than the total score.


Timeline of Intelligence Psychodiagnostics Evolution

Year / Era Key Figures Method Main Shift
1905 Binet, Simon Binet–Simon scale Mental age for schools
1912 Stern IQ formula Quotient as ratio of ages
1916 Terman Stanford–Binet Large-scale US standardization
1917 US military psychologists Army Alpha / Beta Mass group testing
1938 Raven Progressive Matrices Nonverbal focus on fluid g
1939+ Wechsler WAIS / WISC Deviation IQ (M=100, SD=15)
Late 20th – 21st Modern editions Updated norms, computerization Adaptive testing, online formats

Ratio IQ vs Deviation IQ: What's the Difference in Plain Terms

Ratio IQ (Stern) Deviation IQ (Wechsler onward)
Meaning ratio of "mental" and chronological age position relative to peer norm
Mean roughly around 100 when ages match fixed at 100 by standardization
Adults formula breaks down works through age-based norms
Interpretation "how many years ahead/behind" "how many SD from the mean"

Example deviations:

  • 100 — near average;
  • 115 — roughly +1 SD (above average);
  • 85 — roughly −1 SD (below average);
  • extreme tails are rare and require especially careful interpretation.

<WARNING_BLOCK title="Don't Confuse "Online Score" with Clinical Diagnosis">A short web practice is useful for training and orientation. A proper clinical/educational assessment is a standardized procedure with norms, professional qualifications, and ethics. NeuroLab is a training and educational environment, not a replacement for official diagnostics.


What Good IQ Tests Usually Measure (and What They Don't)

Typically Load On

  • Fluid reasoning — reasoning on novel material;
  • Crystallized knowledge (in verbal batteries) — vocabulary, information;
  • Working memory — holding and updating information;
  • Processing speed — rate of simple cognitive operations;
  • Visual-spatial analysis — patterns, assembly, orientation.

Usually Don't Measure Directly

  • moral qualities and character;
  • creativity in all its forms;
  • motivation, sleep, anxiety, and health at test time;
  • narrow specialized professional skill;
  • "life success" as a single quantity.

50–80%
A rough order of what share of school performance variance is often linked to cognitive tests in research — notable, but far from everything. The rest: motivation, quality of teaching, environment, health, opportunity.


Cultural Fairness: The Myth of a "Completely Neutral" Test

Raven's matrices are often called "culture-free." It is more accurate to say: less dependent on language and school knowledge. But factors remain:

  • familiarity with the test format ("test-wiseness");
  • quality of instructions and a calm environment;
  • vision, fatigue, time anxiety;
  • experience with abstract patterns (education, games, STEM).

So an honest history of IQ includes critique as well: tests improved fairness relative to pure teacher subjectivity, but were never a magic mirror of "pure brain outside culture."

💡 Practical Takeaway
Compare yourself with yourself under similar conditions: sleep, time of day, quiet, clear instructions. A single "record for the screenshot" says almost nothing about the trend.


How to Read a Result: A Short Common-Sense Guide

  1. Look at the range and repetition, not a single attempt.
  2. Factor in your state: sleep deprivation easily "eats" tens of minutes of concentration.
  3. If there is a subtest profile — look for strengths, not just the "hole."
  4. Use the result as a learning navigator: where to train reasoning, memory, speed, attention.
  5. Don't turn the score into an identity ("I = a number").

Mini Checklist Before a Serious Attempt

  • slept as well as possible;
  • notifications off;
  • understood the instructions before starting;
  • not taking the test "angrily" after a hard day;
  • ready to accept the result as a starting point, not a verdict.

Connection to Other Cognitive Domains

The history of IQ testing logically aligns with what NeuroLab trains alongside:

Domain Connection to IQ History Why Practice Alongside
Matrix practice the Raven / fluid g line rule extraction
Reaction processing speed less "fog" at the start
Schulte / attention search control resilience to visual noise
Working memory part of Wechsler batteries holding task conditions

Complement your matrix practice with a reaction test: this shows whether overall processing speed is dragging your result down on tired days.
Take the Test Now →

If you often lose your place in a matrix pattern, add short Schulte sessions: this trains visual scanning discipline.
Take the Test Now →


Common Myths About the History and Meaning of IQ

"Binet wanted to rank people permanently."
No: the original commission was about pedagogical help for children with learning difficulties.

"IQ = genetics, period."
Heritability of cognitive measures is researched, but environment, education, health, and opportunity strongly influence the realization of potential. It is not "either/or."

"Raven completely eliminates culture."
Reduces language load — yes. Eliminates culture — no.

"Wechsler just renamed Binet."
No: the key shift is deviation IQ and the expanded profile approach.

"An online test equals a clinical WAIS."
Formats may be related in task concept, but norms, length, proctoring, and purpose differ.

⚠️ Ethics Matter More Than a Pretty Number
The darkest chapters in testing history are tied not to the math of SD=15, but to how results were used against people. Keep the instrument in the zone of learning, self-knowledge, and careful interpretation.


What Changed in the 21st Century

Modern intelligence assessment increasingly includes:

  • computerized and adaptive testing (difficulty adjusts to responses);
  • updating norms for new generations;
  • interest in ability profiles, not just a single score;
  • combining classic tasks with metrics of attention, speed, and stability.

Online platforms like NeuroLab stand in this line as an accessible practice ground: you can regularly train reasoning and related skills, see dynamics, and not wait for "one fateful commission."


Brief Summary

The history of IQ is a journey from the Paris schoolroom Binet–Simon scale to Stern's formula, mass military testing, Raven's nonverbal matrices, and Wechsler's deviation scales.

Remember the framework:

  1. Binet — help with learning through mental age;
  2. Stern — the IQ quotient as a ratio of ages;
  3. Terman / Army — scaling and standardization (with ethical lessons);
  4. Raven — visual reasoning with less language load;
  5. Wechsler — comparison against peer norms and profiling;
  6. Today — practice, repeated measurement, and sober interpretation matter more than the cult of a single number.

If you want not only to know the history but to feel the logic of the tasks yourself — start with the NeuroLab matrix module and look at the trend across several attempts, not one random run.