Skip to content

Course home

Validity

Validity

What you'll learn

  • What validity means and how it differs from reliability.
  • The four AQA validity types: face, concurrent, ecological and temporal validity.
  • How psychologists measure validity in different methods of investigation.
  • How to improve validity while balancing control, realism and ethics.

The starting point: are we measuring the right thing?

Psychology often studies things we cannot directly see, such as anxiety, memory, attachment or obedience. These are called constructs: abstract psychological ideas inferred from behaviour, scores or self-report.

To study a construct, researchers must operationalise it. This means defining it in measurable terms. For example, “memory” might be operationalised as “the number of words correctly recalled from a 20-word list after two minutes”.

Definition

Validity

Validity is the extent to which a test, questionnaire, observation or whole study measures what it claims to measure, and supports the conclusion the researcher wants to draw.

A method can be reliable without being valid. Reliability means consistency: the measure gives similar results over time, across items, or between observers. But a consistently wrong measure is still wrong.

Example

Separating reliability from validity

A researcher measures “exam stress” by asking only: “How many hours did you revise yesterday?”

  1. The measure could be reliable because the same student might give the same number if asked again.

  2. But revision hours do not necessarily measure stress. A calm, organised student may revise for many hours, while a very stressed student may avoid revision.

  3. So the measure has weak validity: it consistently measures revision time, not necessarily exam stress.

Key Idea

Validity is about accuracy of meaning

Reliability asks, “Is the measurement consistent?” Validity asks, “Is this really measuring the intended construct or supporting the intended conclusion?”

The diagram below summarises the four named types of validity and common ways researchers measure or improve them.

Concept map showing face, concurrent, ecological and temporal validity, plus measurement and improvement methods

Validity of measures and validity of studies

Validity can apply to a measurement tool, such as a questionnaire, interview schedule, observation checklist or cognitive test.

It can also apply to a whole investigation. For example, an experiment may validly show that the independent variable caused a change in the dependent variable, or it may be weakened by confounding variables.

Two useful background terms are:

  • Internal validity: whether the study really shows what happened inside the study, such as whether the IV caused the DV.
  • External validity: whether findings generalise beyond the study, such as to other settings, times, people or tasks.

Ecological and temporal validity are both types of external validity.

Face validity

Face validity means a measure appears, on the surface, to measure what it claims to measure.

A questionnaire asking “I feel tense in social situations” looks more face valid for social anxiety than a questionnaire asking only “I enjoy watching television”.

Face validity is usually assessed by inspection. Researchers, experts, teachers, clinicians or sometimes participants look at the measure and judge whether it seems appropriate.

AO3: evaluating face validity

A strength is that it is quick and useful during pilot work. It can identify obviously irrelevant or confusing items before a study begins.

A limitation is that it is subjective and shallow. A measure may look valid but still fail to capture the construct properly. Also, if the aim is too obvious, participants may show demand characteristics, where they change behaviour because they guess what the researcher expects.

Concurrent validity

Concurrent validity means a new measure is compared with an already established measure of the same construct, taken at roughly the same time.

For example, a new depression questionnaire could be compared with a widely used clinical depression inventory. If people who score high on the established test also score high on the new test, this supports concurrent validity.

The established measure is called a criterion: a benchmark used for comparison.

Measuring concurrent validity statistically

Concurrent validity is often assessed using a correlation. A correlation shows the strength and direction of the relationship between two co-variables.

If the data are interval or ratio and suitable for a parametric test, researchers may use Pearson’s r. If the data are ordinal, ranked, or unsuitable for Pearson’s r, they may use Spearman’s rho.

Example

Checking a new anxiety questionnaire

A psychologist gives 20 students a new anxiety questionnaire and an established anxiety scale on the same day.

  1. The design is correlational because the researcher is checking the relationship between two co-variables: scores on the new scale and scores on the established scale.

  2. If the questionnaire scores are interval-level and meet the assumptions for a parametric test, the researcher could use Pearson’s r. If the data are ranked or non-normal, Spearman’s rho would be safer.

  3. The researcher looks up the critical value for the chosen test using N=20N=20N=20, the chosen significance level p<0.05p < 0.05p<0.05, and whether the hypothesis is one-tailed or two-tailed.

  4. If the observed correlation, for example r=+0.78r=+0.78r=+0.78, is larger than the critical value, the researcher rejects H0H_0H0​ and concludes there is evidence for concurrent validity.

  5. If the observed correlation is weak or below the critical value, the researcher retains H0H_0H0​; the new questionnaire does not yet have strong evidence of concurrent validity.

Common Mistake

Significant is not the same as valid

A statistically significant correlation can support concurrent validity, but it does not prove the measure is perfect. The criterion may be flawed, or the two tests may share the same bias, such as both relying on self-report.

Ecological validity

Ecological validity is the extent to which findings can generalise to real-life settings and everyday behaviour.

A laboratory study may have high control but lower ecological validity if the task is artificial. For example, Loftus and Palmer’s 1974 eyewitness testimony study used filmed car accidents rather than real accidents. This helped control the stimulus material, but the task may not fully represent the emotion and consequences of witnessing a genuine crash.

A related term is mundane realism, meaning how much the task resembles real life. Mundane realism can support ecological validity, but they are not exactly the same thing: ecological validity is about generalising findings, not simply where the study took place.

Common Mistake

Ecological validity is not just location

Do not write “field study = ecologically valid” automatically. A study in a natural setting can still use an artificial task, and a lab study can sometimes model a real psychological process well.

Temporal validity

Temporal validity is the extent to which findings remain true across historical time.

Some findings may be time-bound because society changes. For example, Asch’s 1956 conformity research may not produce identical levels of conformity today because education, individualism, gender roles and attitudes to authority have changed. Later research such as Perrin and Spencer’s 1980 study found much lower conformity among engineering students.

However, older research is not automatically invalid. Some psychological processes may be stable, and temporal validity is assessed through replication over time.

When discussing classic studies, remember ethics too. Asch used deception about the true aim of the line-judgement task, so full informed consent was not possible; debriefing was therefore important.

Measuring validity across methods

Different methods have different validity issues.

MethodValidity issueHow validity can be assessed
QuestionnaireDo the items really measure the construct?Face validity, pilot studies, concurrent validity with an established scale
InterviewAre answers honest and relevant?Careful question design, reduced social desirability, comparison with other evidence
ObservationDo behavioural categories represent the target behaviour?Clear operationalisation, training observers, checking whether categories match real behaviour
ExperimentDid the IV really cause the DV?Control of extraneous variables, standardised procedures, random allocation or counterbalancing
CorrelationAre both co-variables validly measured?Use validated measures; remember correlation does not show causation
Case studyDoes rich detail reflect the wider issue?Triangulation using multiple sources, but generalisation may remain limited
Tip

Use triangulation

Triangulation means using more than one method or data source, such as questionnaires, interviews and observations. If they point to the same conclusion, validity is strengthened.

Improving validity

Researchers can improve validity by targeting the specific threat.

For face validity, they can ask experts to review items, remove unclear questions and run a pilot study.

For concurrent validity, they can compare the new measure with an established measure and revise items that do not correlate well.

For ecological validity, they can use more realistic tasks, field settings, naturalistic observation or real-world outcome measures.

For temporal validity, they can replicate the study with contemporary samples, update outdated materials and compare findings across cohorts.

For internal validity, they can control extraneous variables, standardise instructions, use random allocation, counterbalance repeated measures and reduce investigator effects.

Example

Improving a low-validity study

A researcher studies “aggression” by asking students one question: “Are you an aggressive person?”

  1. The construct is poorly operationalised because one direct self-report item may measure self-image more than actual aggression.

  2. To improve measurement validity, the researcher could use a validated aggression questionnaire and compare it with an established scale to assess concurrent validity.

  3. To reduce social desirability, responses could be anonymous and include indirect items rather than only obvious questions.

  4. To improve ecological validity, the researcher could add observation of behaviour in a realistic competitive task, while ensuring protection from harm and the right to withdraw.

  5. To improve temporal validity, the measure could be re-tested with different year groups or future cohorts to check whether the findings still hold.

Validity and ethics

Sometimes improving validity creates ethical pressure. Deception can reduce demand characteristics, but it weakens informed consent and requires careful debriefing. Naturalistic observation can improve ecological validity, but researchers must consider privacy, consent and confidentiality.

A strong exam answer recognises this balance: the most valid method is not always the most ethical, and the most ethical method may sometimes be less natural or less revealing.

Exam technique

In the exam

  1. Name the exact type of validity: face, concurrent, ecological or temporal.

  2. Link it directly to the scenario: say what is being measured, compared, generalised or replicated.

  3. Add AO3 balance: explain one strength and one limitation of that type of validity evidence.

  4. When suggesting improvements, be specific: “use an established anxiety scale” is better than “make it more valid”.

Self review

Check yourself

  • How is concurrent validity different from face validity?
  • Why might a highly controlled lab experiment have lower ecological validity?
  • What could a researcher do to improve the temporal validity of an older study?
PreviousNext

How was this guide?

Teach Genie

Review Validity (A-level only) by teaching Genie

Teach it back in your own words, spot gaps, and remember it better.

Start teaching
Genie and Baby Genie

Lesson

Recap your knowledge with an interactive lesson

8 minute activity

Start lesson

Concept map showing validity split into measurement validity and generalisation validity, with face, concurrent, ecological and temporal validity labelled

Validity is the extent to which a method, measure or study actually measures what it claims to measure and allows accurate conclusions. In psychology this matters because researchers study constructs such as anxiety, obedience and memory, which must be operationalised into something measurable.

Reliability is about consistency, but validity is about truthfulness. A questionnaire could give the same score every week and still be invalid if it is really measuring social desirability rather than genuine anxiety.

AQA often groups the named types into measurement validity and generalisation validity. Face validity and concurrent validity focus on the measure itself, while ecological validity and temporal validity focus on whether findings generalise beyond the original study.

Flashcards

Remember key concepts with flashcards

24 flashcards

Practice flashcards

An abstract psychological idea inferred from behaviour is a [     ].

Validity (A-level only) Revision Guide

  1. A Level
  2. /Psychology
  3. /Validity (A-level only)

Revision guides