What you'll learn
- What validity means, and how it differs from reliability.
- The difference between internal validity and external validity.
- How issues such as researcher bias, demand characteristics and social desirability reduce validity.
- How psychologists improve and assess validity using methods such as blinding, standardisation, and concurrent, predictive, face, content and construct validity checks.
The big idea: what is validity?
In psychology, validity is about truthfulness and accuracy. A study, test, questionnaire or observation is valid if it really investigates what it claims to investigate.
A memory test should measure memory, not reading speed. A stress questionnaire should measure stress, not just how willing someone is to admit they are struggling.
Validity
Validity is the extent to which a method, measure, study or conclusion accurately reflects what it is intended to measure or test.
Validity is different from reliability, which means consistency. A method can be reliable but not valid: it can give the same result repeatedly, while still measuring the wrong thing.

Reliable but not valid
-
A researcher says their questionnaire measures depression, so the intended construct is depression. A construct is a psychological concept that cannot be directly seen, such as anxiety, attachment or intelligence.
-
The questionnaire gives very similar scores when the same participants complete it two weeks apart, so it has good reliability.
-
However, most of the questions are about exam pressure and workload, so the measure may actually be assessing academic stress rather than depression.
-
The conclusion is that the questionnaire may be reliable, but it has poor validity because it does not accurately measure the claimed construct.
Internal validity: is the cause really the cause?
Internal validity is about what happens inside the study. It asks whether the researcher can confidently say that changes in the dependent variable were caused by the independent variable, rather than by some other factor.
Internal validity
Internal validity is the extent to which a study establishes a genuine cause-and-effect relationship between variables, without interference from uncontrolled factors.
To understand internal validity, you need two important terms:
- The independent variable is the variable the researcher manipulates or compares.
- The dependent variable is the outcome that is measured.
- An extraneous variable is any other variable that might affect the dependent variable.
- A confounding variable is an extraneous variable that changes systematically with the independent variable, making it difficult to know what caused the effect.
For example, in Loftus and Palmer’s study of eyewitness testimony (1974), the wording of a question was manipulated. The study had fairly strong internal validity because participants watched the same film clips and the wording was carefully controlled. This helps isolate the effect of the verb used in the question.
Spotting a confounding variable
-
A researcher tests whether caffeine improves memory. The independent variable is caffeine condition, and the dependent variable is number of words recalled.
-
The caffeine group is tested at 9 am, while the no-caffeine group is tested at 4 pm. Time of day now varies systematically with the caffeine condition.
-
Time of day could affect alertness and memory performance, so it is a confounding variable.
-
The study has reduced internal validity because better recall might be due to caffeine, time of day, or both.
-
A better design would test both groups at similar times, or randomly allocate participants across time slots and caffeine conditions.
Internal validity takeaway
High internal validity means you can be more confident that the study’s result was caused by the variable being investigated, not by accidental or uncontrolled influences.
External validity: can the findings generalise?
External validity is about whether findings apply beyond the exact study.
External validity
External validity is the extent to which findings can be generalised to other people, settings, times or situations.
There are three common forms you should know.
Population validity
Population validity concerns whether findings generalise from the sample to the wider target population.
For example, if a study only uses psychology students, it may not generalise well to children, older adults, people from different cultures, or clinical populations.
Ecological validity
Ecological validity concerns whether the findings generalise to real-life settings and behaviours.
A laboratory study may have strong control but feel artificial. For example, watching a filmed car crash in Loftus and Palmer (1974) is not the same as witnessing a real crash, where emotion, surprise and personal consequences may affect memory.
Temporal validity
Temporal validity concerns whether findings remain true over time.
A study from decades ago may be affected by social change. For example, obedience, gender roles, attitudes to mental health, or technology use may differ across historical periods.
Laboratory does not automatically mean invalid
Do not write “this was a lab study, so it has no validity.” A laboratory study may have high internal validity because of strong control, but lower ecological validity because the setting is artificial. Always specify which type of validity you mean.
Specific validity issues
Researcher bias
Researcher bias occurs when a researcher’s expectations influence the procedure, interpretation or recording of data.
This might happen if an observer knows which participants are expected to behave aggressively and therefore notices aggressive acts more readily. In a study such as Raine et al. (1997), which used PET scans with murderers pleading not guilty by reason of insanity, objective brain-imaging data helped reduce some forms of researcher bias, although interpretation of findings still required caution.
Ways to reduce researcher bias include:
- Standardised procedures, where every participant receives the same instructions and conditions.
- Blind techniques, where the participant or researcher does not know key information about the condition.
- Inter-rater reliability checks, where multiple observers independently code behaviour and compare agreement.
Demand characteristics
Demand characteristics are cues in a study that allow participants to guess the aim and change their behaviour.
Participants may try to please the researcher, sabotage the study, or behave in a way they think is expected. This reduces validity because the behaviour is no longer natural.
Ways to reduce demand characteristics include:
- Using a single-blind procedure, where participants do not know the full aim or condition.
- Using filler tasks or less obvious measures.
- Giving standardised instructions.
- Debriefing participants afterwards if deception was used.
Social desirability
Social desirability occurs when participants give answers that make them look good rather than answers that are fully honest.
This is especially likely in research on aggression, prejudice, addiction, criminal behaviour, stress, or mental health. A participant might under-report alcohol use or deny socially unacceptable attitudes.
Ways to reduce social desirability include:
- Anonymous questionnaires.
- Confidential data handling.
- Indirect questioning.
- Carefully worded, non-judgemental items.
- Behavioural measures alongside self-report.
Choosing fixes for a sensitive questionnaire
-
A researcher wants to study bullying behaviour using a self-report questionnaire. The main validity threat is social desirability because participants may not want to admit bullying others.
-
Making the questionnaire anonymous reduces fear of judgement and makes honest answers more likely.
-
Wording items neutrally, such as “How often have you joined in teasing?” rather than “Are you a bully?”, reduces defensiveness.
-
Adding peer nominations or teacher observations could provide triangulation, meaning evidence from more than one method.
-
The study’s validity improves because the researcher is less dependent on one socially sensitive self-report measure.
Ways of dealing with validity problems
A good methods answer does not just identify a validity issue; it explains how to reduce it.
Improve internal validity
You can improve internal validity by:
- Clearly operationalising variables, meaning defining them in observable and measurable terms.
- Using standardised instructions and procedures.
- Controlling extraneous variables.
- Randomly allocating participants to conditions.
- Using matched pairs or repeated measures where appropriate.
- Counterbalancing condition order in repeated measures designs.
- Using single-blind or double-blind procedures.
A double-blind procedure means neither the participant nor the researcher collecting data knows which condition the participant is in. This is especially useful when expectations could influence behaviour or interpretation.
Improve external validity
You can improve external validity by:
- Using a more representative sample.
- Testing in naturalistic or field settings.
- Replicating the study with different ages, cultures or groups.
- Repeating research over time to check temporal validity.
There is often a trade-off: more control can improve internal validity, while more realism can improve ecological validity.
Ethics and validity
Validity improvements must still follow the BPS Code of Ethics and Conduct. Psychologists should consider consent, deception, right to withdraw, protection from harm, confidentiality and debriefing.
Sometimes researchers use deception to reduce demand characteristics, but deception must be justified, not excessive, and followed by a full debrief. In animal research, additional issues include whether animal use is necessary, minimising suffering, appropriate housing and applying the principles of replacement, reduction and refinement.
Do not sacrifice ethics for validity
A study is not “better” simply because deception or stress makes behaviour more realistic. Validity must be balanced against protection from harm, informed consent and the participant’s right to withdraw.
Assessing validity
Psychologists use different types of validity checks depending on what they are evaluating.
Face validity
Face validity means the measure appears, on the surface, to measure what it claims.
A stress questionnaire asking about tension, sleep and worry has higher face validity than one asking only about favourite colours.
Face validity is useful but weak on its own because something can look valid without actually being valid.
Content validity
Content validity means the measure covers the full range of the topic.
For example, a depression scale should include emotional, cognitive, behavioural and physical symptoms rather than only sadness.
Construct validity
Construct validity means the measure genuinely represents the psychological construct it claims to measure.
This is important because constructs like intelligence, attachment or anxiety are theoretical. You cannot directly observe them, so the operationalisation must be justified.
Concurrent validity
Concurrent validity means a new measure is compared with an already established measure taken at the same time.
If a new anxiety scale strongly agrees with a well-established anxiety scale, it has better concurrent validity.
Predictive validity
Predictive validity means a measure accurately predicts future behaviour or outcomes.
For example, if a stress scale predicts later absence from school or work, it has predictive validity.
Assessing a new social anxiety scale
-
To check face validity, look at whether the items seem relevant to social anxiety, such as fear of speaking in groups or avoiding social events.
-
To check content validity, compare the items with expert descriptions of social anxiety and see whether cognitive, emotional, physical and behavioural symptoms are all represented.
-
To check concurrent validity, give participants the new scale and an established social anxiety scale at the same time, then compare the scores.
-
To check predictive validity, test whether high scores predict later avoidance of presentations, social events or interviews.
-
To check construct validity, see whether scores relate to similar constructs, such as general anxiety, but are not simply measuring unrelated traits.
Validity and statistical conclusions
Validity is not only about the design; it also affects conclusions. In Component 2, choosing the correct statistical test supports the validity of your inference.
For Eduqas, you need to choose tests using the design, aim and level of measurement: Spearman’s rho for correlations, chi-square for nominal associations or differences, the binomial sign test for related nominal data, Mann-Whitney U for unrelated non-parametric comparisons, Wilcoxon signed-ranks for related non-parametric comparisons, and related or unrelated t-tests for interval data meeting parametric assumptions.
At the usual significance level of p≤0.05p \le 0.05p≤0.05, you compare an observed value with a critical value from a table, using the correct rule for that test and whether the hypothesis is one-tailed or two-tailed.
Statistics cannot rescue poor validity
A significant result does not prove that the study measured the right thing. If a questionnaire has poor construct validity or an experiment has a major confounding variable, the conclusion may still be weak even if the statistical test is correct.
In the exam
-
Name the specific type of validity you are discussing: internal, external, ecological, population, temporal, face, content, construct, concurrent or predictive.
-
Link the validity issue directly to the scenario, study or practical investigation rather than giving a generic definition.
-
For AO3, explain the consequence: does the issue weaken cause-and-effect, generalisation, measurement accuracy, or confidence in the conclusion?
Check yourself
- What is the difference between internal validity and external validity?
- How could demand characteristics and social desirability affect a questionnaire study differently?
- Which type of validity would you use to see whether a new test predicts future behaviour?
