What you'll learn
- Why Gould argued that early IQ testing confused intelligence with culture, language and education.
- The AO1 “story” of the study: background, method, results and conclusions.
- How to evaluate Gould using validity, reliability, ethics, sampling bias and ethnocentrism.
- How to compare Gould with the paired contemporary study, Hancock et al. (2011).
Why this study matters
Gould (1982) is the classic core study for the OCR H567 key theme Measuring differences in the Individual Differences area.
The big question is: when psychologists measure a difference between people, are they measuring a real psychological difference — or are they measuring bias in the tool?
The main takeaway
Gould’s argument is that IQ tests can look scientific and objective, but if the test favours one language, culture or education level, the score may reflect test bias rather than genuine intelligence.
Key background concepts
Intelligence and IQ
Intelligence is usually defined as the ability to learn, reason, solve problems and adapt to new situations. However, psychologists disagree about whether intelligence is one general ability or many different abilities.
An IQ test is a standardised test designed to produce a numerical score for intellectual ability. In early testing, this score was often treated as if it measured a fixed, inherited trait.
IQ
IQ, or intelligence quotient, is a score intended to represent intellectual ability compared with a reference group. The problem is that the score depends heavily on what the test includes and who it was standardised on.
Mental age
Early intelligence testing used the idea of mental age. For example, if a child performed like the average 10-year-old, they might be said to have a mental age of 10.
This idea becomes much more problematic when applied to adults. Calling an adult “mentally aged 13” can be misleading because adult intelligence, education, language experience and cultural knowledge do not develop in a simple age-line.
Test bias
Test bias
Test bias occurs when a test systematically advantages or disadvantages certain groups for reasons unrelated to the ability being measured.
For Gould, the key issue was that the Army IQ tests used during the First World War claimed to measure innate intelligence, but many items also required English language, schooling, test-taking experience and knowledge of American culture.
Spotting cultural bias in a test item
Suppose an intelligence test asks: “Crisco is a: medicine, disinfectant, toothpaste, or food product.”
- The test claims to measure reasoning or intelligence, so the relevant skill should be problem-solving rather than product knowledge.
- The item also requires familiarity with an American brand name and English vocabulary, which are extra demands unrelated to intelligence.
- A recent immigrant might answer incorrectly because they do not know the brand, not because they lack reasoning ability.
- Therefore, the item risks underestimating the intelligence of people outside the dominant culture.
Historical background: Binet, Yerkes and eugenics
The original intelligence tests were developed by Alfred Binet in France to identify children who needed educational support. Binet did not intend IQ to be used as a fixed label of worth.
In the USA, psychologists such as Robert Yerkes adapted intelligence testing for large-scale use. During the First World War, Yerkes and colleagues tested huge numbers of army recruits using group intelligence tests.
This happened in a wider social context of eugenics.
Eugenics
Eugenics is the belief that society should be “improved” by encouraging some groups to reproduce and discouraging or preventing others. It is now recognised as deeply unethical and discriminatory.
Gould argued that IQ testing was used to support racist and anti-immigration views, especially the idea that some ethnic groups were naturally less intelligent.
Historical language
Gould discusses offensive historical labels such as “moron” and “feeble-minded”. These terms are included only because they were used in the historical material; they are not acceptable modern psychological language.
The study story: Gould (1982)
Aim
Gould aimed to challenge the claim that the US Army intelligence tests measured innate intelligence fairly. He wanted to show that the tests were biased and that the results had been misused to support discriminatory conclusions.
Design
Gould’s work was not a laboratory experiment. It was a historical review and secondary data analysis.
Secondary data analysis
Secondary data analysis means analysing data that were originally collected by someone else, often for a different purpose.
Gould examined Yerkes’ published army testing data, test materials, administration procedures and the conclusions drawn from the results.
Sample and source material
The original data came from around 1.75 million US Army recruits tested during 1917–1918. They were male military recruits, including many immigrants and men with different levels of education and English language ability.
Gould’s “sample” was therefore not a new group of participants. It was the existing historical record of Army Alpha and Army Beta testing.
Materials and apparatus
The main testing materials were:
- Army Alpha: a written verbal test for literate English-speaking recruits.
- Army Beta: a pictorial or non-verbal test intended for illiterate recruits or those who did not speak English well.
- Individual examinations: used for some recruits who failed earlier testing.

Procedure
Recruits were usually given Army Alpha if they were judged literate in English. Those who failed Alpha, or who were considered illiterate or non-English-speaking, could be moved to Army Beta. Some who failed Beta received individual assessment.
Gould reviewed the content of the tests and the testing conditions. He considered whether low scores could be explained by factors such as poor English, limited schooling, unfamiliarity with American culture, rushed group testing and anxiety in a military setting.
Results
Yerkes’ interpretation of the results was alarming: the average white American recruit was said to have a mental age of about 13. This supported the famous phrase “a nation of morons”.
The tests also appeared to show that some immigrant groups scored lower than others, and that Black recruits scored particularly low. Hereditarian psychologists used this to argue that intelligence differences were inherited and linked to race or nationality.
Gould disagreed. He argued that the pattern made much more sense if the scores reflected acculturation, education and English language exposure. Immigrants who had lived in the USA longer tended to do better, which suggests the test measured familiarity with American culture rather than fixed innate intelligence.
Acculturation
Acculturation means adapting to the language, customs and knowledge of a new culture.
Conclusions
Gould concluded that the Army tests lacked validity as measures of innate intelligence. They were affected by language, culture, schooling and administration conditions.
He also argued that the results were socially dangerous because they were used to justify eugenics and immigration restriction.
Evaluating Gould: AO3
Validity
The key evaluation point is construct validity.
Construct validity
Construct validity is whether a test really measures the psychological concept it claims to measure.
Yerkes claimed the tests measured innate intelligence. Gould argued they actually measured a mixture of intelligence, English literacy, cultural knowledge, education and confidence with formal tests.
This is a strong criticism because if the measure is invalid, the impressive number of participants does not rescue the conclusion.
Big sample does not guarantee good measurement
A huge sample can make findings look powerful, but if the test itself is biased, the data may be consistently wrong on a large scale.
Reliability
The Army tests were standardised in the sense that many recruits received similar materials. However, testing conditions were not always consistent. Recruits could be tested in large groups, under time pressure, in unfamiliar military settings.
Gould’s own review is also open to the criticism of researcher bias. He selected and interpreted historical evidence to support an argument. However, because he used published materials and quoted test items, other researchers can inspect the same sources.
Sampling bias and generalisability
The original Army sample was large, but not representative of the whole population. It included military-age men, not women, children, older adults or people outside the recruitment system.
There were also major differences in education, social class, immigration history and language background. Gould’s point is that these differences were treated as evidence of innate intelligence when they may have reflected social opportunity.
Ethics and socially sensitive research
The original Army testing raises serious ethical issues. Recruits were in a military hierarchy, so genuine consent and right to withdraw were limited. The results labelled groups as inferior and contributed to discriminatory policies.
Under the modern BPS Code of Human Research Ethics, psychologists should consider respect, responsibility, integrity and social impact. Gould’s study is important because it exposes how psychological measurement can harm people when used carelessly.
Socially sensitive research
Research on intelligence, race, class or criminality is socially sensitive because findings can influence public policy and how groups are treated.
Applying Gould to a new situation: AO2
Gould is very useful for applying to modern examples, such as schools, workplaces or immigration systems using ability tests.
Applying Gould to a modern testing scenario
A school uses an English-only reasoning test to place newly arrived pupils into ability groups.
- The school claims the test measures reasoning ability, so you should ask whether the score reflects reasoning or another factor.
- Newly arrived pupils may have weaker English or less familiarity with British school-style tests, which are confounding variables.
- If these pupils score lower, Gould would suggest this may show language and cultural bias rather than lower intelligence.
- A fairer approach would use multiple measures, language-appropriate assessment and caution when interpreting scores.
Comparison with Hancock et al. (2011)
The paired contemporary study is Hancock et al. (2011), often known as Hungry like the wolf. It investigated whether psychopathic murderers used language differently from non-psychopathic murderers.
| Feature | Gould (1982) | Hancock et al. (2011) |
|---|---|---|
| Key theme | Measuring differences | Measuring differences |
| Difference measured | Intelligence/IQ | Psychopathy and language use |
| Method | Historical review and secondary data analysis | Interviews and computerised language analysis |
| Sample | Historical data from US Army recruits | Male murderers in Canadian prisons |
| Main issue | IQ tests were culturally biased | Whether language patterns reveal psychopathic traits |
| Ethics | Social harm from racist and anti-immigration interpretations | Risks of labelling offenders as psychopaths |
Both studies show that measuring individual differences can be useful, but also risky. Gould shows how measurement can be distorted by culture and politics. Hancock et al. used more modern, systematic tools, but still raises questions about validity, sampling and socially sensitive labels.
Links to the Individual Differences area
The Individual Differences area studies how people vary, for example in intelligence, personality, mental health or criminal behaviour.
Gould fits this area because he focuses on whether psychologists can fairly measure differences between people. His study also links strongly to debates:
- Nature/nurture: Yerkes emphasised inherited intelligence; Gould emphasised environment, education and culture.
- Reductionism: IQ testing reduced complex human ability to a single score.
- Ethnocentrism: the tests favoured American English-speaking cultural knowledge.
- Usefulness: Gould’s critique helps improve fair testing, but the original testing was harmful.
- Psychology as a science: numbers and tests can look objective, but scientific measurement depends on valid concepts and fair methods.
In the exam
- For AO1, make clear that Gould analysed Yerkes’ Army Alpha and Beta testing; Gould did not personally test the recruits.
- For AO3, prioritise validity, cultural bias, sampling bias, ethics and socially sensitive research rather than giving vague strengths and weaknesses.
- When comparing with Hancock et al. (2011), focus on the shared theme of measuring differences, then contrast method, sample, validity and ethical risks.
Check yourself
- Why did Gould argue that Army Alpha and Army Beta did not validly measure innate intelligence?
- How could language, education and acculturation affect IQ test scores?
- What is one similarity and one difference between Gould (1982) and Hancock et al. (2011)?
