What you'll learn
- What reliability means in psychological research and why it matters.
- The difference between internal reliability and external reliability.
- How to assess reliability using split-half, test-retest and inter-rater methods.
- How to improve reliability in your own Component 2 practical work.
AO1: What is reliability?
In psychology, researchers often measure things that cannot be seen directly, such as anxiety, aggression, memory or wellbeing. These are called constructs: psychological qualities or processes that we infer from behaviour, responses or scores.
A measure is useful only if it gives results that are consistent. If the same questionnaire, observation schedule or test produces wildly different results for no good reason, it is hard to trust the findings.
Reliability
Reliability means the consistency of a measure, procedure or set of observations. A reliable measure should produce similar results when the same thing is being measured in the same way.
Reliability is closely linked to the scientific goal of replicability, which means that findings can be repeated using the same method. If a study cannot be repeated consistently, psychologists should be cautious about drawing firm conclusions from it.

Reliability is not the same as validity
Validity
Validity means whether a test, observation or study measures what it claims to measure.
A measure can be reliable but not valid. For example, a questionnaire might consistently give the same score, but if the questions do not really measure anxiety, it is not valid.
Reliable but wrong
A broken set of scales that always adds two kilograms is reliable because it is consistent, but it is not valid because it does not give the true weight.
Reliability does not prove validity
Do not write that a study is “valid because it is reliable”. Reliability helps research quality, but a consistent measure can still consistently measure the wrong thing.
Internal reliability
Internal reliability
Internal reliability means consistency within a measure, especially when a questionnaire, scale or test has several items that are supposed to measure the same construct.
For example, if you create a ten-item questionnaire to measure self-esteem, all ten items should be pulling in roughly the same direction. If half the questions seem to measure confidence, while the other half measure popularity, the scale may have poor internal reliability.
Assessing internal reliability: split-half reliability
Split-half reliability
Split-half reliability assesses internal reliability by dividing a test or questionnaire into two comparable halves, scoring each half separately, and checking whether the two halves produce similar results.
A common method is to split the questionnaire into odd-numbered and even-numbered items. The researcher then compares participants’ scores on the two halves. If the two sets of scores are strongly positively correlated, the measure has good internal reliability.
A correlation coefficient is a number showing the strength and direction of a relationship between two sets of scores. It is often written as rrr, where values closer to plus one show a stronger positive relationship.
Using split-half reliability
A researcher creates a 12-item stress questionnaire and splits it into odd-numbered and even-numbered items. The correlation between the two half-scores is r=+0.82r = +0.82r=+0.82.
- Identify the type of reliability being tested: the researcher is checking whether items inside the same questionnaire are consistent, so this is internal reliability.
- Interpret the coefficient: r=+0.82r = +0.82r=+0.82 is a strong positive correlation, meaning participants who scored highly on one half also tended to score highly on the other half.
- Make a judgement: the questionnaire appears to have good split-half reliability, although the researcher should still check whether the items are valid measures of stress.
External reliability
External reliability
External reliability means consistency beyond one single measurement, such as consistency over time, across different researchers, or across different situations.
External reliability is especially important in observations, experiments and repeated testing. If findings depend too much on the exact observer, day or situation, the research may be unreliable.
Assessing external reliability: test-retest reliability
Test-retest reliability
Test-retest reliability assesses whether the same test gives similar results when the same participants complete it again after a time gap.
The researcher gives participants the same measure twice, then compares Time 1 and Time 2 scores. If scores are very similar, the measure has good test-retest reliability.
This is most appropriate when the construct should be fairly stable, such as personality traits or long-term attitudes. It is less useful for states that naturally change quickly, such as mood, tiredness or stress before an exam.
Interpreting test-retest evidence
A researcher gives a personality questionnaire to 30 students in September and again two weeks later. The correlation between the two sets of scores is r=+0.89r = +0.89r=+0.89.
- Decide whether test-retest is suitable: personality is expected to be relatively stable over two weeks, so repeating the questionnaire is appropriate.
- Interpret the relationship: r=+0.89r = +0.89r=+0.89 shows that students’ scores were very similar across the two testing occasions.
- Draw a conclusion: the questionnaire has strong external reliability over time, assuming the second test was not affected by memory of the first test.
Time gaps matter
If the retest happens too soon, participants may remember their original answers. If it happens too late, the construct may genuinely have changed. Either problem can distort test-retest reliability.
Assessing external reliability: inter-rater reliability
Inter-rater reliability
Inter-rater reliability means the extent to which two or more observers, coders or raters agree when recording the same behaviour or material.
This is especially important in observations and content analysis. For example, if two observers are recording aggression in a playground, they should agree on what counts as “aggression”. If one observer counts teasing as aggression and the other does not, the data will be inconsistent.
Researchers can assess inter-rater reliability using percentage agreement or a correlation. In simple observational coding, percentage agreement is often used:
Percentage agreement=number of agreementsnumber of agreements+number of disagreements×100\text{Percentage agreement} = \frac{\text{number of agreements}}{\text{number of agreements} + \text{number of disagreements}} \times 100Percentage agreement=number of agreements+number of disagreementsnumber of agreements×100Calculating inter-rater agreement
Two observers watch the same video and code 50 behaviours. They agree on 42 behaviours and disagree on 8.
- Work out the total number of coding decisions: 42+8=5042 + 8 = 5042+8=50.
- Substitute into the formula: 4250×100=84\frac{42}{50} \times 100 = 845042×100=84.
- Interpret the result: 84 percent agreement suggests acceptable inter-rater reliability, although the observers should review the disagreements to improve the coding system.
A useful benchmark
In many classroom examples, around 80 percent agreement is treated as a reasonable minimum for inter-rater reliability. In real research, the acceptable level depends on the task and consequences of error.
AO2: Applying reliability to your own investigation
In Component 2, reliability is often applied to a novel scenario or your own practical work. You should be able to say exactly what could go wrong and how to check or improve it.
If you use a questionnaire
You can improve reliability by:
- using clear, unambiguous questions;
- avoiding double-barrelled questions, such as “I feel anxious and unhappy”;
- using consistent response options, such as the same five-point rating scale throughout;
- assessing split-half reliability if several items measure the same construct.
If you use an observation
You can improve reliability by:
- creating operationalised behavioural categories, meaning categories that define behaviours clearly and objectively;
- making categories mutually exclusive, so one behaviour cannot fit two categories at once;
- training observers using the same coding manual;
- checking inter-rater reliability before analysing the final data.
Operationalisation
Operationalisation means defining a variable or behaviour in a precise, measurable way so that researchers know exactly what to record.
If you repeat a measure
You can improve reliability by:
- keeping instructions and materials the same each time;
- using the same testing conditions where possible;
- choosing a sensible time gap;
- checking test-retest reliability if the construct should remain stable.
Choosing a reliability check
A student investigates whether teenagers show more helping behaviour when alone or in groups. Two observers use a checklist to record behaviours such as “picks up dropped item” and “offers verbal help”.
- Identify the method: this is an observation using behavioural categories, so observer judgement could affect the results.
- Choose the reliability assessment: inter-rater reliability is appropriate because two observers are coding the same behaviours.
- Improve the procedure: the student should operationalise each helping behaviour, train both observers using the same examples, and calculate percentage agreement before the main study.
Ways of dealing with reliability issues
A strong answer should go beyond saying “make it reliable”. You need to explain the specific method.
Improving reliability
Reliability is improved by reducing random variation in how data are collected, scored and interpreted.
Useful strategies include:
- Standardised procedures: give every participant the same instructions, materials and time limits.
- Pilot studies: run a small trial first to detect confusing questions, unclear categories or timing problems.
- Observer training: practise coding examples until observers apply categories consistently.
- Clear scoring rules: decide in advance how responses will be marked or coded.
- Equipment checks: make sure apparatus, online surveys or recording devices work consistently.
- Reviewing weak items: remove or rewrite questionnaire items that do not match the rest of the scale.
Repeating is not always fixing
Repeating a test can reveal poor reliability, but it does not automatically solve it. You must identify the source of inconsistency, such as unclear questions or poorly trained observers.
Ethics and reliability
Reliability checks still need to follow the BPS Code of Ethics and Conduct. This includes informed consent, protection from harm, right to withdraw, confidentiality and debriefing.
For example, if you video-record behaviour so that two observers can code it, participants should normally know they are being recorded, recordings should be stored securely, and identities should be protected.
If animal research is involved, additional ethical issues apply, including minimising suffering, using the fewest animals necessary, providing suitable housing and considering alternatives where possible.
AO3: Evaluating reliability
Reliability is a major strength because it increases confidence that findings are not just due to chance, unclear procedures or researcher inconsistency. Reliable methods are easier to replicate, which supports psychology’s scientific status.
However, reliability has limits. A measure can be highly reliable but invalid. Very rigid standardisation may also reduce ecological validity if the situation becomes artificial. Inter-rater reliability can be high because categories are simple, but this may miss subtle behaviour. Test-retest reliability can be affected by memory, practice effects or genuine change over time.
You can also link reliability to the wider statistics content. Reliability checks often use descriptive comparisons, percentages or correlations. In inferential testing, Eduqas expects you to choose tests using the aim, design and level of measurement: for example, Spearman’s rho for correlations, chi-square for nominal frequency data, Wilcoxon or Mann-Whitney U for ordinal differences, and related or unrelated t-tests for interval differences. The usual significance convention is p≤0.05p \leq 0.05p≤0.05, but remember: statistical significance is not the same thing as reliability.
In the exam
- Define reliability as consistency, then specify the type: internal, external, split-half, test-retest or inter-rater.
- Apply your answer to the scenario by naming the actual measure, observer, questionnaire, time gap or behavioural category.
- For AO3, explain both sides: reliability improves trust and replicability, but it does not guarantee validity.
Check yourself
- What is the difference between internal reliability and external reliability?
- Which reliability check would you use for an observation with two coders?
- Why can a measure be reliable but still not valid?
