What you'll learn
- What reliability means in psychological research.
- How reliability applies across experiments, observations, questionnaires, interviews, correlations, case studies and content analysis.
- How to assess reliability using test-retest and inter-observer methods.
- How to improve reliability without confusing it with validity.
The basic idea: reliability means consistency
In research methods, a method of investigation is the way psychologists collect data: for example, an experiment, observation, questionnaire, interview, case study, correlation or content analysis.
Reliability is about whether that method produces consistent results. If the same thing is measured again under similar conditions, a reliable method should give similar findings.
Reliability
Reliability means consistency: a research method, measure or observation is reliable if it produces similar results when repeated under similar conditions, or when different researchers use it in the same way.
A bathroom scale gives a simple analogy. If it gives wildly different readings every time you step on it, it is unreliable. In psychology, the same issue applies to a depression questionnaire, an observation of aggression, or a content analysis of media messages.
Here is the big picture of the two main AQA reliability checks and how researchers improve them.

Reliability is not the same as validity
A method can be reliable but not valid.
Validity
Validity means accuracy: whether a study or measure really measures what it claims to measure.
For example, a questionnaire might give very consistent scores each time it is completed, so it has good reliability. But if it is supposed to measure anxiety and actually measures social desirability — how much someone wants to look good to others — then it has poor validity.
Reliability versus validity
Reliability asks, “Is it consistent?” Validity asks, “Is it measuring the right thing?” A measure can be reliable without being valid, but it is hard for a measure to be valid if it is very unreliable.
Reliability across methods of investigation
Reliability matters in all psychological methods, but it looks slightly different depending on the method.
Experiments
In an experiment, the researcher manipulates an independent variable, which is the factor being changed, and measures a dependent variable, which is the outcome.
Reliability improves when the procedure is standardised, meaning every participant is treated in the same way as far as possible. This includes the same instructions, same timing, same materials and same scoring rules.
For example, Loftus and Palmer’s 1974 study on leading questions used controlled film clips and standardised questions. This helps reliability because another researcher could repeat the procedure more easily. However, reliability does not automatically mean the task is realistic, so validity still needs evaluating.
Observations
In an observation, researchers record behaviour as it happens. Reliability depends heavily on whether observers agree about what they saw.
If one observer records “aggression” but another does not, the problem may be that the behavioural categories are unclear. Researchers improve reliability by making categories precise, such as “hits another child with hand or object” rather than “acts aggressively”.
Questionnaires and interviews
A questionnaire is a set of written questions. An interview is a spoken research conversation.
Structured questionnaires and structured interviews are usually more reliable because everyone receives the same questions in the same order. Unstructured interviews may be less reliable because the interviewer can ask different follow-up questions each time.
Correlations
A correlation investigates whether two co-variables are related. Reliability matters because both variables must be measured consistently. If one measure is unreliable, the correlation may be weaker or misleading.
For example, if a stress scale gives unstable scores from one day to the next for no clear reason, it may be hard to judge whether stress is genuinely related to sleep.
Case studies and content analysis
A case study is an in-depth investigation of one individual, group or institution. Case studies can be difficult to replicate exactly, so reliability may be lower.
A content analysis is a method for systematically coding communication, such as adverts, newspapers, films or interview transcripts. Reliability improves when coders use clear categories and independently code the same material.
Thinking reliability only applies to questionnaires
Reliability applies to every research method. In an observation it may mean observer agreement; in an experiment it may mean standardised procedures; in content analysis it may mean consistent coding.
Measuring reliability: test-retest reliability
Test-retest reliability
Test-retest reliability is assessed by giving the same test, questionnaire or measure to the same people on two different occasions, then comparing the two sets of scores.
The logic is simple: if the measure is reliable, participants should get similar scores both times, assuming the thing being measured has not genuinely changed.
Researchers often compare the two sets of scores using a correlation coefficient, which is a number showing the strength and direction of a relationship between two variables. A strong positive correlation suggests good test-retest reliability.
In AQA Psychology, you do not usually need to calculate this from raw data, but you should understand the decision: if scores at Time 1 and Time 2 are very similar, reliability is high.
Assessing test-retest reliability
A researcher creates a 30-item questionnaire to measure trait anxiety. The same 40 students complete it in September and again two weeks later. The researcher obtains r=0.86r = 0.86r=0.86.
-
Identify the reliability method: the same questionnaire has been given to the same participants on two occasions, so this is test-retest reliability.
-
Choose the appropriate statistical idea: because the researcher is comparing two sets of questionnaire scores from the same people, a correlation such as Pearson’s r may be suitable if the scores are treated as interval data and assumptions are met. If scores were ordinal or not normally distributed, Spearman’s rho may be more suitable.
-
Interpret the relationship: r=0.86r = 0.86r=0.86 is a strong positive correlation, so students who scored high the first time tended to score high the second time.
-
Make the reliability judgement: the questionnaire appears to have good test-retest reliability, although the researcher should consider whether anxiety genuinely changed during the two-week interval.
-
If testing significance formally, the researcher would look up the critical value for the chosen correlation test using the sample size and significance level, commonly p<0.05p < 0.05p<0.05. If the calculated value is equal to or greater than the critical value, the null hypothesis of no relationship would be rejected.
Choosing the time gap
The time gap in test-retest reliability needs to be sensible.
If the gap is too short, participants may remember their previous answers. This can create artificially high reliability due to practice effects or memory.
If the gap is too long, the psychological characteristic may genuinely change. For example, stress may change across an exam period, and mood can change from day to day.
Test-retest is not suitable for everything
Test-retest reliability works best for fairly stable characteristics, such as personality traits. It is less useful for temporary states, such as current mood, because real change can be mistaken for unreliability.
Measuring reliability: inter-observer reliability
Inter-observer reliability
Inter-observer reliability is the extent to which two or more observers produce the same record of the same behaviour or material.
This is especially important in observations and content analysis. It can also apply when researchers code interview transcripts.
The usual procedure is:
- Create clear behavioural categories or coding categories.
- Train two or more observers to use the same categories.
- Ask them to record the same behaviour or code the same material independently.
- Compare their records to see how much they agree.
A simple method is percentage agreement, which calculates the percentage of observations where the observers gave the same code.
percentage agreement=number of agreementstotal observations×100\text{percentage agreement} = \frac{\text{number of agreements}}{\text{total observations}} \times 100percentage agreement=total observationsnumber of agreements×100Calculating percentage agreement
Two observers watch the same playground video and code behaviour in 50 time intervals. They agree in 42 intervals.
-
Identify what counts as agreement: an agreement occurs when both observers record the same behavioural category for the same interval.
-
Substitute into the percentage agreement formula:
-
Compare with a practical criterion: many researchers would see 84% agreement as reasonably good, especially if the categories are complex.
-
Improve before the main study: the observers should review the 8 disagreements, clarify category definitions and practise again to increase reliability.
More advanced research may use Cohen’s kappa, a statistic that adjusts for agreement that could occur by chance. For A-Level Psychology, it is usually enough to know that high agreement between observers suggests good inter-observer reliability.
Assuming two observers always solve the problem
Two observers can still share the same bias, especially if they both know the aim of the study. Inter-observer reliability is stronger when observers are trained, independent and ideally blind to the hypothesis.
Improving reliability
Reliability is not just measured after the study. Good researchers build it into the design from the start.
Operationalise variables clearly
Operationalisation
Operationalisation means defining a variable precisely so it can be measured or manipulated.
“Anxiety” is vague. “Score on a 20-item anxiety questionnaire” is operationalised. “Aggression” is vague. “Number of times a child hits, kicks or pushes another child in 10 minutes” is clearer.
Clear operationalisation makes it more likely that another researcher could repeat the study in the same way.
Standardise the procedure
Standardisation
Standardisation means keeping the procedure the same for all participants, such as the same instructions, timings, materials, environment and scoring method.
Standardisation reduces the influence of extraneous variables, which are unwanted factors that might affect the results. For example, if some participants receive extra encouragement and others do not, differences in scores may be due to the researcher rather than the variable being studied.
Train observers and interviewers
Observers should practise using the same coding system before the main study. Interviewers should use an agreed interview schedule if reliability is important.
This reduces observer bias, which occurs when an observer’s expectations influence what they record.
Use pilot studies
A pilot study is a small trial run before the main investigation. It helps researchers spot unclear questions, confusing instructions, overlapping categories or practical problems.
For example, a pilot observation might show that “physical aggression” and “rough play” are too difficult to separate. The researcher can then refine the categories before collecting the main data.
Use clear scoring rules
Questionnaires, interviews and content analyses are more reliable when researchers use a scoring manual. This explains exactly how answers or behaviours should be coded.
However, there is a trade-off. Closed questions and rigid coding can improve reliability, but they may reduce richness and validity if they oversimplify complex experiences.
Easy exam phrase
To explain how to improve reliability, write: “The researcher should operationalise the variables, standardise the procedure and use trained independent observers so that the study can be repeated consistently.”
AO3 evaluation: why reliability matters
Strength: supports scientific replication
A major strength of reliability is that it supports replication, which means repeating a study to see whether similar findings are obtained. Psychology aims to be scientific, so reliable methods help researchers check whether findings are dependable rather than accidental.
Strength: improves confidence in findings
If a questionnaire gives similar scores across time, or observers agree closely, researchers can be more confident that the data reflect a real pattern. Low reliability adds “noise” to the data and can hide genuine effects.
Limitation: reliability does not prove validity
A reliable measure may consistently measure the wrong thing. For example, a memory test might produce consistent scores because participants learn the task format, not because it accurately measures everyday memory.
This is important in essay evaluation: never treat reliability as the only mark of good research.
Limitation: high standardisation may reduce realism
Highly standardised procedures can improve reliability, but they may reduce ecological validity, which is the extent to which findings reflect real-life behaviour. For example, a tightly controlled laboratory task may be easy to replicate but unlike real social interaction.
Ethical angle
Improving reliability should not come at the expense of ethics. Researchers still need appropriate consent, avoidance of unnecessary deception, the right to withdraw, protection from harm, confidentiality and a proper debrief.
This is especially relevant in observations. Covert observation may improve natural behaviour and reliability, but it can raise consent and privacy issues.
In the exam
-
Identify the method first: for a questionnaire or test, think test-retest; for observations or coding, think inter-observer reliability.
-
Link your answer to the scenario: name exactly what would be repeated, who would repeat it, and what scores or records would be compared.
-
Evaluate carefully: reliability is a strength if findings are consistent and replicable, but it does not automatically prove validity.
Check yourself
- How would you test the reliability of a new depression questionnaire?
- Why might two observers disagree when recording the same behaviour?
- Give one way to improve reliability in an observation and one way to improve it in an interview.
