Skip to content
MathsGenie logo
Open app

Course home

  1. A Level
  2. Psychology Eduqas
  3. Revision guides

Reliability

What you'll learn

  • What reliability means in psychological research and why it matters.
  • The difference between internal reliability and external reliability.
  • How to assess reliability using split-half, test-retest and inter-rater methods.
  • How to improve reliability in your own Component 2 practical work.

AO1: What is reliability?

In psychology, researchers often measure things that cannot be seen directly, such as anxiety, aggression, memory or wellbeing. These are called constructs: psychological qualities or processes that we infer from behaviour, responses or scores.

A measure is useful only if it gives results that are consistent. If the same questionnaire, observation schedule or test produces wildly different results for no good reason, it is hard to trust the findings.

Definition

Reliability

Reliability means the consistency of a measure, procedure or set of observations. A reliable measure should produce similar results when the same thing is being measured in the same way.

Reliability is closely linked to the scientific goal of replicability, which means that findings can be repeated using the same method. If a study cannot be repeated consistently, psychologists should be cautious about drawing firm conclusions from it.

Concept map showing internal reliability, external reliability and ways to improve reliability

Reliability is not the same as validity

Definition

Validity

Validity means whether a test, observation or study measures what it claims to measure.

A measure can be reliable but not valid. For example, a questionnaire might consistently give the same score, but if the questions do not really measure anxiety, it is not valid.

Analogy

Reliable but wrong

A broken set of scales that always adds two kilograms is reliable because it is consistent, but it is not valid because it does not give the true weight.

Common Mistake

Reliability does not prove validity

Do not write that a study is “valid because it is reliable”. Reliability helps research quality, but a consistent measure can still consistently measure the wrong thing.

Internal reliability

Definition

Internal reliability

Internal reliability means consistency within a measure, especially when a questionnaire, scale or test has several items that are supposed to measure the same construct.

For example, if you create a ten-item questionnaire to measure self-esteem, all ten items should be pulling in roughly the same direction. If half the questions seem to measure confidence, while the other half measure popularity, the scale may have poor internal reliability.

Assessing internal reliability: split-half reliability

Definition

Split-half reliability

Split-half reliability assesses internal reliability by dividing a test or questionnaire into two comparable halves, scoring each half separately, and checking whether the two halves produce similar results.

A common method is to split the questionnaire into odd-numbered and even-numbered items. The researcher then compares participants’ scores on the two halves. If the two sets of scores are strongly positively correlated, the measure has good internal reliability.

A correlation coefficient is a number showing the strength and direction of a relationship between two sets of scores. It is often written as rrr, where values closer to plus one show a stronger positive relationship.

Example

Using split-half reliability

A researcher creates a 12-item stress questionnaire and splits it into odd-numbered and even-numbered items. The correlation between the two half-scores is r=+0.82r = +0.82r=+0.82.

  1. Identify the type of reliability being tested: the researcher is checking whether items inside the same questionnaire are consistent, so this is internal reliability.
  2. Interpret the coefficient: r=+0.82r = +0.82r=+0.82 is a strong positive correlation, meaning participants who scored highly on one half also tended to score highly on the other half.
  3. Make a judgement: the questionnaire appears to have good split-half reliability, although the researcher should still check whether the items are valid measures of stress.

External reliability

Definition

External reliability

External reliability means consistency beyond one single measurement, such as consistency over time, across different researchers, or across different situations.

External reliability is especially important in observations, experiments and repeated testing. If findings depend too much on the exact observer, day or situation, the research may be unreliable.

Assessing external reliability: test-retest reliability

Definition

Test-retest reliability

Test-retest reliability assesses whether the same test gives similar results when the same participants complete it again after a time gap.

The researcher gives participants the same measure twice, then compares Time 1 and Time 2 scores. If scores are very similar, the measure has good test-retest reliability.

This is most appropriate when the construct should be fairly stable, such as personality traits or long-term attitudes. It is less useful for states that naturally change quickly, such as mood, tiredness or stress before an exam.

Example

Interpreting test-retest evidence

A researcher gives a personality questionnaire to 30 students in September and again two weeks later. The correlation between the two sets of scores is r=+0.89r = +0.89r=+0.89.

  1. Decide whether test-retest is suitable: personality is expected to be relatively stable over two weeks, so repeating the questionnaire is appropriate.
  2. Interpret the relationship: r=+0.89r = +0.89r=+0.89 shows that students’ scores were very similar across the two testing occasions.
  3. Draw a conclusion: the questionnaire has strong external reliability over time, assuming the second test was not affected by memory of the first test.
Common Mistake

Time gaps matter

If the retest happens too soon, participants may remember their original answers. If it happens too late, the construct may genuinely have changed. Either problem can distort test-retest reliability.

Assessing external reliability: inter-rater reliability

Definition

Inter-rater reliability

Inter-rater reliability means the extent to which two or more observers, coders or raters agree when recording the same behaviour or material.

This is especially important in observations and content analysis. For example, if two observers are recording aggression in a playground, they should agree on what counts as “aggression”. If one observer counts teasing as aggression and the other does not, the data will be inconsistent.

Researchers can assess inter-rater reliability using percentage agreement or a correlation. In simple observational coding, percentage agreement is often used:

Percentage agreement=number of agreementsnumber of agreements+number of disagreements×100\text{Percentage agreement} = \frac{\text{number of agreements}}{\text{number of agreements} + \text{number of disagreements}} \times 100Percentage agreement=number of agreements+number of disagreementsnumber of agreements​×100
Example

Calculating inter-rater agreement

Two observers watch the same video and code 50 behaviours. They agree on 42 behaviours and disagree on 8.

  1. Work out the total number of coding decisions: 42+8=5042 + 8 = 5042+8=50.
  2. Substitute into the formula: 4250×100=84\frac{42}{50} \times 100 = 845042​×100=84.
  3. Interpret the result: 84 percent agreement suggests acceptable inter-rater reliability, although the observers should review the disagreements to improve the coding system.
Tip

A useful benchmark

In many classroom examples, around 80 percent agreement is treated as a reasonable minimum for inter-rater reliability. In real research, the acceptable level depends on the task and consequences of error.

AO2: Applying reliability to your own investigation

In Component 2, reliability is often applied to a novel scenario or your own practical work. You should be able to say exactly what could go wrong and how to check or improve it.

If you use a questionnaire

You can improve reliability by:

  • using clear, unambiguous questions;
  • avoiding double-barrelled questions, such as “I feel anxious and unhappy”;
  • using consistent response options, such as the same five-point rating scale throughout;
  • assessing split-half reliability if several items measure the same construct.

If you use an observation

You can improve reliability by:

  • creating operationalised behavioural categories, meaning categories that define behaviours clearly and objectively;
  • making categories mutually exclusive, so one behaviour cannot fit two categories at once;
  • training observers using the same coding manual;
  • checking inter-rater reliability before analysing the final data.
Definition

Operationalisation

Operationalisation means defining a variable or behaviour in a precise, measurable way so that researchers know exactly what to record.

If you repeat a measure

You can improve reliability by:

  • keeping instructions and materials the same each time;
  • using the same testing conditions where possible;
  • choosing a sensible time gap;
  • checking test-retest reliability if the construct should remain stable.
Example

Choosing a reliability check

A student investigates whether teenagers show more helping behaviour when alone or in groups. Two observers use a checklist to record behaviours such as “picks up dropped item” and “offers verbal help”.

  1. Identify the method: this is an observation using behavioural categories, so observer judgement could affect the results.
  2. Choose the reliability assessment: inter-rater reliability is appropriate because two observers are coding the same behaviours.
  3. Improve the procedure: the student should operationalise each helping behaviour, train both observers using the same examples, and calculate percentage agreement before the main study.

Ways of dealing with reliability issues

A strong answer should go beyond saying “make it reliable”. You need to explain the specific method.

Key Idea

Improving reliability

Reliability is improved by reducing random variation in how data are collected, scored and interpreted.

Useful strategies include:

  • Standardised procedures: give every participant the same instructions, materials and time limits.
  • Pilot studies: run a small trial first to detect confusing questions, unclear categories or timing problems.
  • Observer training: practise coding examples until observers apply categories consistently.
  • Clear scoring rules: decide in advance how responses will be marked or coded.
  • Equipment checks: make sure apparatus, online surveys or recording devices work consistently.
  • Reviewing weak items: remove or rewrite questionnaire items that do not match the rest of the scale.
Common Mistake

Repeating is not always fixing

Repeating a test can reveal poor reliability, but it does not automatically solve it. You must identify the source of inconsistency, such as unclear questions or poorly trained observers.

Ethics and reliability

Reliability checks still need to follow the BPS Code of Ethics and Conduct. This includes informed consent, protection from harm, right to withdraw, confidentiality and debriefing.

For example, if you video-record behaviour so that two observers can code it, participants should normally know they are being recorded, recordings should be stored securely, and identities should be protected.

If animal research is involved, additional ethical issues apply, including minimising suffering, using the fewest animals necessary, providing suitable housing and considering alternatives where possible.

AO3: Evaluating reliability

Reliability is a major strength because it increases confidence that findings are not just due to chance, unclear procedures or researcher inconsistency. Reliable methods are easier to replicate, which supports psychology’s scientific status.

However, reliability has limits. A measure can be highly reliable but invalid. Very rigid standardisation may also reduce ecological validity if the situation becomes artificial. Inter-rater reliability can be high because categories are simple, but this may miss subtle behaviour. Test-retest reliability can be affected by memory, practice effects or genuine change over time.

You can also link reliability to the wider statistics content. Reliability checks often use descriptive comparisons, percentages or correlations. In inferential testing, Eduqas expects you to choose tests using the aim, design and level of measurement: for example, Spearman’s rho for correlations, chi-square for nominal frequency data, Wilcoxon or Mann-Whitney U for ordinal differences, and related or unrelated t-tests for interval differences. The usual significance convention is p≤0.05p \leq 0.05p≤0.05, but remember: statistical significance is not the same thing as reliability.

Exam technique

In the exam

  1. Define reliability as consistency, then specify the type: internal, external, split-half, test-retest or inter-rater.
  2. Apply your answer to the scenario by naming the actual measure, observer, questionnaire, time gap or behavioural category.
  3. For AO3, explain both sides: reliability improves trust and replicability, but it does not guarantee validity.
Self review

Check yourself

  • What is the difference between internal reliability and external reliability?
  • Which reliability check would you use for an observation with two coders?
  • Why can a measure be reliable but still not valid?
PreviousNext

How was this guide?

Teach Genie

Review Reliability by teaching Genie

Teach it back in your own words, spot gaps, and remember it better.

Start teaching
Genie and Baby Genie

Lesson

Recap your knowledge with an interactive lesson

8 minute activity

Start lesson

Concept map of reliability showing internal reliability with split-half, external reliability with test-retest and inter-rater, and a note that reliability does not always mean validity

Psychologists often study constructs such as anxiety, memory, or aggression. A measure is only useful if it gives consistent results rather than changing randomly.

Reliability means the consistency of a measure, procedure, or set of observations. If the same thing is measured in the same way, the results should be similar.

Reliability supports replicability, but it is not the same as validity. A broken scale that always adds 2 kilograms is reliable because it is consistent, but not valid because it is wrong.

Flashcards

Remember key concepts with flashcards

24 flashcards

Practice flashcards

A psychological [     ] is inferred from behaviour, responses or scores because it cannot be seen directly.

Reliability Revision Guide

  1. A Level
  2. /Psychology
  3. /Reliability

Revision notes for Eduqas A Level Psychology Reliability. Open the guide for explanations and worked examples. Written against the Eduqas A Level Psychology (A290QS) specification, so the content matches what's examinable rather than general Psychology background.

Revision guides