
4. Evaluation of Kohlberg (1968) (AO3)
Evaluating developmental research requires focusing on how well the theory stands up to different populations, genders, and time spans.
Methodological Strengths
- Longitudinal Design: Following the same individuals over 12 years eliminates individual differences (participant variables) as confounding variables. We can be highly confident that changes in moral reasoning were due to development rather than differences in upbringing or personality.
- Cross-Cultural Support: By studying boys in various cultures (e.g., Taiwan, Mexico), Kohlberg could argue that his developmental stages were a universal feature of human cognitive growth, not just an artifact of American culture.
Methodological Weaknesses
- Gender Bias (Androcentrism): The original sample was entirely male. Carol Gilligan (1982) famously criticized Kohlberg, arguing that females tend to develop a morality of care (prioritizing relationships) rather than a morality of justice (prioritizing abstract rules). Therefore, Kohlberg’s stages are biased against females, often categorizing them at a lower moral stage (Stage 3) than males (Stage 4).
- Hypothetical Dilemmas (Low Ecological Validity): Asking teenagers what Heinz should do is very different from seeing how they would act in a real-world moral crisis. The choices lack real-world consequences, meaning the study may measure verbal reasoning rather than actual behavior.
- Attrition: In longitudinal studies, participants frequently drop out over time (due to moving away, losing interest, etc.). This can leave a biased sub-sample at the end of the study.
5. Mathematical Application: Analyzing Research Data
In Component 2, you are expected to handle quantitative data, recognize levels of measurement, and make statistical decisions. Let's apply these skills to a practical scenario based on these studies.
Suppose a contemporary researcher wants to compare the moral stage of two groups of adolescents (14-year-olds vs 18-year-olds) using Kohlberg’s categorization. To do this, the researcher must understand what kind of data they are gathering.
Levels of Measurement
- Nominal Data: Data categorized into distinct, named groups with no inherent numerical value or order (e.g., Male/Female, Obedient/Disobedient).
- Ordinal Data: Data that can be ranked or ordered, but where the intervals between the ranks are not equal or mathematically measurable (e.g., Kohlberg's moral stages 1 to 6).
- Interval Data: Data measured on a continuous scale with equal, precisely defined intervals between units (e.g., temperature in Celsius, time in seconds).
Classifying levels of measurement and selecting a statistical test
Imagine a researcher wants to compare the moral reasoning of a group of 14-year-old boys and a group of 18-year-old boys. She interviews them and places each boy into one of Kohlberg's stages (Stage 1 to 6).
How should the researcher classify this data, and which inferential statistical test should she use to see if the older group has significantly higher moral reasoning?
-
Identify the level of measurement: The data consists of assigning participants to Stage 1, Stage 2, Stage 3, up to Stage 6. Although these are numbered, the distance between Stage 1 and Stage 2 is not mathematically identical to the distance between Stage 5 and Stage 6. Therefore, the data can be ordered but not precisely measured. This means the data is ordinal.
-
Identify the experimental design: The researcher is comparing two distinct, unrelated groups of participants (14-year-olds vs 18-year-olds). Because a participant can only belong to one age group, this is an independent groups design (unrelated design).
-
Select the correct inferential test: Looking at our criteria for choosing a statistical test:
- We are looking for a difference between two groups.
- The experimental design is unrelated (independent groups).
- The level of measurement is ordinal.
Using these criteria, the correct statistical test to use is the Mann-Whitney U test.
Choosing your statistical test
Memorize this simple decision hierarchy for differences:
- Unrelated + Nominal →\to→ Chi-Square
- Related + Nominal →\to→ Binomial Sign Test
- Unrelated + Ordinal →\to→ Mann-Whitney U
- Related + Ordinal →\to→ Wilcoxon Signed-Ranks
- Unrelated + Interval →\to→ Unrelated t-test
- Related + Interval →\to→ Related t-test
In the exam
- Differentiate AO1, AO2, and AO3: When asked about these studies, keep descriptions (AO1) concise and focused on the key procedures and numbers (e.g., 65%, 75 boys, 12 years). Use most of your writing time to develop well-structured, evaluative paragraphs (AO3).
- Use the P.E.E.L. structure for evaluation:
- Point: State the strength or weakness (e.g., "One weakness of Kohlberg's study is gender bias").
- Explanation: Explain why this is a weakness in psychological terms (e.g., "He used an all-male sample, but Carol Gilligan argues women use a care-based moral system rather than a justice-based one").
- Evidence/Example: Provide detail (e.g., "Gilligan notes that this frequently results in women being categorized as Stage 3, whereas men are categorized as Stage 4").
- Link: Connect back to the question (e.g., "This reduces the population validity of Kohlberg's findings, meaning they cannot be generalized to females").
- Link to the BPS Code of Ethics: When discussing the ethics of Milgram, always refer directly to British Psychological Society (BPS) guidelines: informed consent, deception, protection from harm, confidentiality, debriefing, and the right to withdraw.
Check yourself
- Why is it technically incorrect to call Milgram's baseline (1963) study a laboratory experiment?
- What are the three levels of moral development according to Kohlberg, and what defines the reasoning at each level?
- If a researcher replicates Milgram's study and records how many volts each participant administers, what level of measurement is the voltage data?
