- How psychologists organise quantitative and qualitative data.
- How to calculate and interpret the mean, median, mode and range.
- How to choose and read frequency tables, bar charts, histograms, line graphs and scatter diagrams.
- How sampling, rounding and graph interpretation affect the quality of conclusions.
In GCSE Psychology, you are not expected to do advanced statistics. You are expected to handle research data sensibly: summarise it, draw graphs, spot patterns and judge whether conclusions are justified.
For AO1, you need the correct terms. For AO2, you apply them to a source or scenario. For AO3, you evaluate whether the data are reliable, valid and useful.
Reliability means consistency: would the measure or procedure produce similar results if repeated? Validity means accuracy: does the study really measure what it claims to measure?
Psychology often uses scores, frequencies, percentages and averages.
A decimal is a number written with a decimal point, such as 0.75. A fraction shows a part of a whole, such as one quarter. A percentage means “out of 100”. A ratio compares amounts, such as 2:3.
A quick percentage formula is:
percentage=partwhole×100\text{percentage}=\frac{\text{part}}{\text{whole}}\times 100percentage=wholepart×100
Estimation means making a sensible approximate calculation to check whether your answer is reasonable. Standard form is a way of writing very large or very small numbers as:
N=A×10nwhere 1≤A<10N=A\times 10^n \quad \text{where } 1 \le A < 10N=A×10nwhere 1≤A<10
You are unlikely to need complex standard form in GCSE Psychology, but you should recognise that it is just a compact way of writing numbers.
Calculating a percentage
- A researcher finds that 12 out of 40 participants reported poor sleep, so the part is 12 and the whole is 40.
- Substitute into the formula: 1240×100=30\frac{12}{40}\times 100=304012×100=30.
- Interpret the result: 30% of the sample reported poor sleep.
Quantitative and qualitative data
Quantitative data are numerical data, such as a memory score out of 20. Qualitative data are descriptive data, such as written comments in an interview.
Quantitative data are useful because they are easy to compare, graph and calculate averages from. However, numbers can miss detail. For example, a wellbeing score tells you how much, but not always why.
Qualitative data give richer detail about thoughts, feelings and experiences. However, they are harder to summarise, and interpretation can be more subjective.
Primary and secondary data
Primary data are collected first-hand by the researcher for the current study. Secondary data already exist, such as government statistics, previous research reports or school attendance records.
Primary data are usually more closely matched to the research aim, but they take time and effort to collect. Secondary data can be quick and useful, but may not have been collected in exactly the way the psychologist needs.
Classifying research data
- A questionnaire score of 18 out of 30 is quantitative because it is numerical.
- A participant’s written explanation of why they felt anxious is qualitative because it uses words and meaning.
- If the psychologist collected both during their own study, they are primary data; if the scores came from a published report, they are secondary data.
A population is the whole group a researcher is interested in. A sample is the smaller group who actually take part.
Sampling affects validity
A sample should represent the target population as closely as possible. If the sample is biased, the researcher’s conclusions may not generalise well to the wider population.
A sampling method is the way participants are selected. A larger sample can reduce the effect of unusual individual scores, but size alone does not guarantee validity. A large biased sample is still biased.
For example, if a sleep study only uses volunteers from one school, it may not represent all teenagers. This affects population validity, which means whether findings can be generalised to the wider group.
Judging a sample
- A researcher wants to study sleep in UK teenagers, so the target population is UK teenagers.
- The sample is 25 volunteers from one school, which may over-represent pupils who are confident, available or interested in sleep.
- A better approach would be to include pupils from several schools and try to sample different ages, genders and backgrounds, improving representativeness.
Bigger does not always mean better
Do not write “the sample is large, so it is automatically valid.” A large sample can still be unrepresentative if it is selected in a biased way.
A measure of central tendency is a way of describing the “typical” score in a set of data.
The arithmetic mean is found by adding all the scores and dividing by the number of scores.
mean=sum of scoresnumber of scores\text{mean}=\frac{\text{sum of scores}}{\text{number of scores}}mean=number of scoressum of scores
The mean uses every score, so it is useful for precise comparisons. However, it can be distorted by an outlier, which is an unusually high or low score.
The median is the middle score when all scores are placed in order. If there are two middle scores, add them together and divide by 2.
The median is useful when there are outliers because it is less affected by extreme scores.
The mode is the most common score or category. In grouped data, the modal class is the group interval with the highest frequency.
The mode is useful for categories, such as the most common type of response, but sometimes there is no mode or more than one mode.
The range is a simple measure of dispersion, which means how spread out the scores are.
range=highest score−lowest score\text{range}=\text{highest score}-\text{lowest score}range=highest score−lowest score
A small range suggests scores are clustered together. A large range suggests scores are more spread out. However, the range is affected strongly by outliers.
Calculating average and spread
- For the scores 4, 6, 6, 7 and 12, add them to find the total: 35.
- Divide by the number of scores: 355=7\frac{35}{5}=7535=7, so the mean is 7.
- Put the scores in order and select the middle value: 4, 6, 6, 7, 12, so the median is 6.
- Identify the most common score: 6 appears twice, so the mode is 6.
- Subtract the lowest score from the highest score: 12−4=812-4=812−4=8, so the range is 8.
Choosing the best average
Use the mean when the data are fairly balanced, the median when there are outliers, and the mode when you need the most common category or score.
Forgetting to order the median
The median only works after the scores have been put in order from lowest to highest.
A significant figure is a digit that contributes to the precision of a number, starting from the first non-zero digit. A decimal place is a digit after the decimal point.
In Psychology, you should use an appropriate number of significant figures. This means giving enough detail without pretending your data are more precise than they really are.
For example, a mean of 17.6666667 from test scores is usually too precise. Reporting 17.7 or 17.67 is clearer, depending on the context.
Rounding a mean score
- A researcher calculates a mean score of 17.6666667 from questionnaire data.
- To round to 3 significant figures, keep 1, 7 and the first 6; the next digit is 6, so round up to 17.7.
- To round to 2 decimal places, keep two digits after the decimal point; 17.6666667 becomes 17.67.
A frequency is how often something occurs. A frequency table organises scores or categories alongside their frequencies. A tally is a quick counting mark used before writing the final frequency.
Frequency tables help you translate raw data into a clearer numerical form. From there, you can create graphs.
A bar chart uses separated bars and is best for categories or discrete scores. A histogram uses touching bars and is used for continuous grouped data, such as hours slept. A line graph plots points and joins them to show change across ordered values, such as changes over time. A pie chart shows parts of a whole, usually as percentages.
This diagram compares common displays you may need to construct or interpret.

Making a frequency table
- For the scores 2, 3, 3, 4, 2, 5 and 3, identify the possible scores: 2, 3, 4 and 5.
- Count each score: 2 occurs twice, 3 occurs three times, 4 occurs once and 5 occurs once.
- The modal score is 3 because it has the highest frequency.
- A bar chart would show separated bars for scores 2, 3, 4 and 5, with the tallest bar at 3.
Histograms use continuous groups
In a histogram, bars touch because the groups are continuous intervals. If the intervals are unequal, comparing bar heights can be misleading.
A scatter diagram plots two variables for each participant or case. Each dot represents one pair of scores.
A correlation is a relationship between two variables.
- A positive correlation means both variables tend to increase together.
- A negative correlation means one variable tends to increase as the other decreases.
- Zero correlation means there is no clear relationship.
Correlation is not causation
A scatter diagram can show that two variables are related, but it does not prove that one variable caused the other.
Identifying a correlation
- A scatter diagram shows hours of social media use on the x-axis and wellbeing score on the y-axis.
- The dots slope downwards from left to right, so higher social media use is associated with lower wellbeing scores.
- This is a negative correlation, but you cannot conclude that social media use directly causes lower wellbeing because other variables may be involved.
A normal distribution is a bell-shaped pattern of scores. Most scores cluster around the centre, and fewer scores appear at the very low and very high ends.
In a normal distribution, the mean, median and mode are all at the centre. The curve is symmetrical, meaning the left and right sides are balanced.

Normal distributions are useful because many human characteristics are roughly spread in this way. However, not all psychological data are normal. If scores are skewed, the mean may be pulled towards extreme values.
Recognising a normal distribution
- A frequency pattern of 1, 3, 7, 10, 7, 3 and 1 has the highest frequency in the middle.
- The frequencies on each side match, so the pattern is symmetrical.
- This suggests an approximately normal distribution, with the mean, median and mode near the centre.
In the exam, you may need to turn a table into a graph, read values from a graph, or explain a trend in words.
When interpreting a graph, focus on:
- The x-axis and y-axis labels.
- The scale used on each axis.
- The highest and lowest values.
- The overall trend or pattern.
- Any unusual scores or outliers.
When plotting two variables, place one variable on each axis, plot each pair of scores accurately, and then interpret the pattern. For experimental data, you might compare condition means using a bar chart. For correlational data, you usually use a scatter diagram.
In the exam
- Match the graph to the data: bar chart for categories, histogram for continuous grouped data, line graph for ordered change, scatter diagram for correlation.
- When calculating averages, show enough working to make your method clear and round to a sensible number of significant figures.
- For AO3, do not just describe the data; judge whether the sample, measure and graph support a valid conclusion.
Check yourself
- When would the median be a better average than the mean?
- What is the difference between a bar chart and a histogram?
- Why can a scatter diagram show a relationship but not prove causation?