What you'll learn
- How criminological psychologists investigate eye-witness testimony, including laboratory experiments, field experiments and case studies.
- How to choose samples, analyse quantitative and qualitative data, and select the right inferential test.
- How to judge research using reliability, validity, objectivity, credibility and ethics.
- How to plan a practical investigation linked to a contemporary key question such as: Is eye-witness testimony too unreliable to trust?
Starting point: what counts as criminological research?
Criminological psychology applies psychological theories and methods to crime-related issues: witnesses, suspects, offenders, victims, juries and the justice system.
Eye-witness testimony
Eye-witness testimony (EWT) is evidence given by someone who saw a crime or important event. It may involve recall of what happened, recognition of a suspect, or identification from a line-up.
A classic study is Loftus and Palmer (1974), a laboratory experiment showing that leading questions can alter participants’ memory for a car crash. A useful real-world contrast is Yuille and Cutshall (1986), who studied witnesses to a real shooting and found their recall was still fairly accurate months later. Together, these studies are excellent for AO3 because they disagree partly due to method and realism.
Research methods for assessing EWT effectiveness
Laboratory experiments
A laboratory experiment is a study in a controlled setting where the researcher manipulates an independent variable (IV) and measures a dependent variable (DV).
In EWT research, the IV might be the type of question asked, such as “smashed” versus “hit”. The DV might be estimated speed, number of accurate details recalled, or whether a participant falsely reports seeing broken glass.
Laboratory experiments are strong for control and replication, but may lack ecological validity, meaning the task may not feel like a real crime situation.
Field experiments
A field experiment takes place in a more natural setting while still manipulating an IV. For example, researchers might stage a theft in a classroom and later test witnesses using a standard interview or cognitive interview.
Field experiments often have better ecological validity than lab experiments, but there is less control over extraneous variables such as where participants were standing, how anxious they felt, or how well they saw the event.
Case studies
A case study is an in-depth investigation of one person, group, event or case. In criminological psychology, this might involve detailed interviews with witnesses from a real crime, case files, CCTV evidence and police records.
Case studies provide rich detail and high real-world relevance, but their findings may not generalise to other crimes or witnesses.
The main trade-off is control versus realism: as research becomes more realistic, it often becomes harder to control.

Control versus realism
Loftus and Palmer (1974) gives strong control over variables, while Yuille and Cutshall (1986) gives stronger ecological validity. In evaluation, explain how the method affects the conclusions we can draw about real witnesses.
Choosing a method for an EWT study
A researcher wants to test whether the cognitive interview improves recall of a staged theft in a classroom.
- The researcher needs to compare two interview conditions, so an experiment is appropriate rather than a purely descriptive case study.
- Because the theft happens in a classroom rather than an artificial lab, the study has features of a field experiment.
- The IV is interview type: cognitive interview or standard interview. The DV could be the number of accurate details recalled.
- A strength is improved ecological validity because participants witness an event in a familiar setting. A weakness is reduced control because some participants may have a better view than others.
Sampling in criminological psychology
A sample is the group of participants selected from a wider target population, which is the group the researcher wants to generalise to.
| Sampling method | What it means | Strength | Weakness |
|---|---|---|---|
| Random | Every member of the target population has an equal chance of selection | Reduces researcher bias | Needs a full sampling frame |
| Stratified | The sample reflects key subgroups in the target population | More representative | Time-consuming to organise |
| Volunteer | Participants self-select, often by responding to an advert | Easy to gather motivated participants | Volunteer bias: unusual participants may come forward |
| Opportunity | Researcher uses whoever is available | Quick and practical | Often unrepresentative |
Random does not mean casual
A random sample is not “whoever happens to be there”. That is opportunity sampling. Random sampling requires a fair selection process, such as a random number generator from a complete list.
Reliability, validity, objectivity and credibility
Reliability
Reliability means consistency. A reliable measure gives similar results across time, researchers or items when nothing important has changed.
In EWT research, reliability might involve two researchers independently scoring recall transcripts and checking whether they agree. This is called inter-rater reliability.
Validity
Validity means the study measures what it claims to measure. In EWT, a valid study should measure genuine witness memory, not just participants’ ability to watch a short film in a classroom.
Important types include internal validity, where the IV really causes changes in the DV, and ecological validity, where findings apply to real-world settings.
Objectivity means findings are not shaped by personal opinion. Researchers can improve objectivity by using standardised instructions, clear scoring systems and blind rating, where scorers do not know which condition a participant was in.
Credibility is especially important in qualitative research. It means the interpretation is believable and well-supported by evidence, for example through careful coding, reflexivity and triangulation.
AO3 evaluation chain
For strong evaluation, use the chain: method feature → effect on evidence → consequence for justice. For example: “A lab study has low ecological validity, so it may underestimate the stress of a real crime, meaning it may not fully predict real witness accuracy.”
Ethics: BPS and HCPC
The British Psychological Society Code of Ethics and Conduct (2009) sets out principles researchers should follow.
Key ethical issues in criminological research include:
- Informed consent: participants should know what they are agreeing to where possible.
- Deception: sometimes used in staged crime studies, but it must be justified.
- Right to withdraw: participants can leave the study and withdraw their data.
- Protection from harm: crime materials, distressing videos or mock interrogations must be risk assessed.
- Confidentiality: personal data and responses should be protected.
- Debrief: participants should be told the true purpose afterwards and given support if needed.
The Health and Care Professions Council (HCPC) is relevant when psychologists undertake formulation or intervention, for example with offenders, victims or witnesses. Practitioners should work within competence, manage risk, communicate clearly, keep records and protect service users.
Crime research can create distress
Even a “mock crime” can upset participants, especially if it involves violence, threat or personal experiences of crime. Ethical risk management is not an add-on; it shapes the design from the start.
Quantitative data analysis
Quantitative data is numerical data, such as number of correct details recalled, confidence ratings, reaction times or identification accuracy.
Measures of central tendency
A measure of central tendency describes the typical score.
- Mean: add all scores and divide by the number of scores.
- Median: the middle score when scores are ordered.
- Mode: the most common score.
Measures of dispersion
A measure of dispersion describes spread.
- Range: highest score minus lowest score.
- Standard deviation: the typical distance of scores from the mean.
Calculating descriptive statistics
A small EWT study records the number of accurate details recalled: 6, 7, 7, 8, 12.
- Calculate the mean by adding the scores and dividing by the number of scores:
- Calculate the range by subtracting the lowest score from the highest score:
- Calculate the standard deviation by comparing each score with the mean, squaring the deviations, finding their average, then taking the square root:
- Substitute and simplify:
- Interpret the result: the mean recall score is 8 details, but the standard deviation of about 2.10 shows that scores vary noticeably around that average.
Frequency tables, bar charts and histograms
A frequency table shows how often each score or category occurs. A bar chart is used for discrete categories, such as interview type. A histogram is used for continuous data grouped into intervals, such as age bands or recall-score ranges.
Correlations
A correlation tests whether two co-variables are related, such as witness confidence and identification accuracy. A positive correlation means both increase together. A negative correlation means one increases as the other decreases. Correlation does not prove causation.
Meta-analysis
A meta-analysis statistically combines findings from multiple studies. For example, Köhnken et al. (1999) used meta-analysis to examine the cognitive interview and found it generally increased correct recall, although it could also increase incorrect details. This is useful because it looks beyond one small study.
Inferential statistics
Inferential statistics help decide whether a pattern in the data is likely to be a real effect or could have occurred by chance.
Statistical significance
A result is statistically significant when the probability of the result occurring by chance is low enough to reject the null hypothesis. In psychology, the default level is usually p≤.05p \le .05p≤.05, though p≤.10p \le .10p≤.10 may be used for more exploratory research and p≤.01p \le .01p≤.01 for a stricter test.
Levels of measurement
You choose a test partly based on the level of data:
- Nominal data: categories, such as guilty/not guilty.
- Ordinal data: ranked or ordered data, such as confidence ratings.
- Interval or ratio data: numerical data with equal intervals, such as recall scores.
Normal distributions are symmetrical bell-shaped distributions. Skewed distributions have a long tail on one side. The Edexcel tests here are mainly non-parametric, meaning they do not require the assumption of a normal distribution.
Use this decision tree to choose the correct test.

Which test when?
- Chi-square: association between categories, using nominal data.
- Spearman’s rho: correlation between two co-variables, usually ordinal or ranked data.
- Mann-Whitney U: difference between two independent groups.
- Wilcoxon signed-ranks: difference between two related conditions, such as repeated measures or matched pairs.
Choosing and interpreting an inferential test
A class investigates whether a cognitive interview improves recall. The same participants first complete a standard interview and later complete a cognitive interview after watching a different but similar crime clip.
- The research question asks about a difference, not a correlation or association.
- The same participants take part in both interview conditions, so the design is repeated measures.
- The correct test is therefore Wilcoxon signed-ranks.
- The hypothesis is directional if the researcher predicts the cognitive interview will produce more accurate details, so a one-tailed test is suitable.
- For Wilcoxon, the observed value must be equal to or less than the critical value to be significant. If the observed value is 6 and the critical value is 8, the result is significant because 6≤86 \le 86≤8.
- The researcher can reject the null hypothesis and conclude that interview type affected recall accuracy at the chosen significance level.
Critical value rule changes by test
For chi-square and Spearman’s rho, the observed value usually needs to be equal to or greater than the critical value. For Mann-Whitney U and Wilcoxon, the observed value usually needs to be equal to or less than the critical value.
One-tailed and two-tailed tests
A one-tailed test is used with a directional hypothesis, such as “the cognitive interview will increase recall”. A two-tailed test is used with a non-directional hypothesis, such as “interview type will affect recall”.
A Type I error is a false positive: rejecting the null hypothesis when it is actually true. A Type II error is a false negative: failing to reject the null hypothesis when there really is an effect. Using p≤.10p \le .10p≤.10 increases the risk of Type I errors, while using p≤.01p \le .01p≤.01 reduces Type I errors but can increase Type II errors.
Qualitative data analysis
Qualitative data is non-numerical data, such as interview answers about why a defendant may have committed a crime.
Thematic analysis involves reading the data, coding meaningful units, grouping codes into themes, reviewing themes, naming them and using extracts as evidence.
Grounded theory goes further: the researcher develops a theory from the data rather than starting with a fixed theory. It uses repeated coding and comparison until a clear explanation emerges.
For example, if participants discuss a defendant in a courtroom drama, themes might include “poverty”, “peer pressure”, “anger” and “lack of impulse control”. These themes could then be converted into frequencies for quantitative analysis.
Key question: is EWT too unreliable to trust?
A balanced answer should not simply say “yes” or “no”. Laboratory research such as Loftus and Palmer (1974) suggests memory is reconstructive and vulnerable to misleading information. However, field and case-based evidence such as Yuille and Cutshall (1986) suggests real witnesses can sometimes be accurate, especially for central details.
A sensible contemporary conclusion is that EWT should not be abolished, but it should be treated carefully: police should use improved interviewing methods, avoid leading questions, record interviews, and support vulnerable witnesses.
Practical investigation: a cognitive interview experiment
A suitable 6.5.1 practical could be an experiment into whether the cognitive interview improves recall of a filmed or staged event.
A clear plan would include:
- Research question: Does the cognitive interview improve recall accuracy compared with a standard interview?
- Hypothesis: Participants will recall more accurate details after a cognitive interview than after a standard interview.
- Method: Experiment using independent groups or repeated measures.
- Sampling: Opportunity or volunteer sampling from classmates, with limitations discussed.
- Ethics: Consent, right to withdraw, confidentiality, debrief and protection from harm.
- Data collection: Standardised crime clip, standardised interview schedule and scoring sheet.
- Data analysis: Mean, median, range, standard deviation and an appropriate inferential test.
- Results and discussion: Decide whether to reject the null hypothesis and link findings to EWT reliability.
- Improvements: Larger sample, better controls, random allocation, blind scoring and more realistic materials.
Making your practical stronger
Use a clear scoring system before collecting data, such as “one mark for each accurate person, action, object or location detail”. This improves objectivity and inter-rater reliability.
In the exam
- For AO1, describe the method or statistic accurately using key terms such as IV, DV, reliability, validity, observed value and critical value.
- For AO2, apply directly to the scenario: identify the sample, design, data type and likely test from the details given.
- For AO3, make a clear judgement about impact: explain how a strength or weakness affects confidence in conclusions about crime, witnesses or justice.
Check yourself
- Which method gives more control: a lab experiment, field experiment or case study?
- Why would Wilcoxon be used for a repeated-measures cognitive interview study?
- How could a researcher make qualitative interview analysis more credible?