- How to tell the difference between quantitative, qualitative, primary and secondary data.
- How to calculate and interpret GCSE descriptive statistics, including averages, range, ratios, fractions and percentages.
- How to choose suitable tables, charts and graphs for psychology data.
- How to judge reliability, validity and bias in research.
In psychology, data means the information researchers collect to answer a research question. Analysing research is the process of organising that information, summarising it, representing it clearly, and judging whether the findings are trustworthy.
Analysis is more than maths
At GCSE, analysing research means both doing calculations and making judgements: What does the pattern show? Is the measure reliable? Is the sample valid? Could bias have affected the results?
Quantitative data
Quantitative data is numerical data, such as scores, ratings, frequencies or percentages. It can be counted, measured and compared using statistics.
Qualitative data
Qualitative data is non-numerical data, usually words, descriptions or explanations. It might come from interviews, open questionnaire answers or observations.
Quantitative data is useful because it is easy to summarise, compare and display in graphs. Qualitative data is useful because it gives richer detail about what people think, feel or experience.
Primary data
Primary data is data collected first-hand by the researcher for their own study.
Secondary data
Secondary data is data that already exists, collected by someone else, such as official statistics, school records or published research.
| Data type | Main strength | Possible weakness |
|---|
| Quantitative | Easy to compare and analyse statistically | May miss detail or personal meaning |
| Qualitative | Rich, detailed and can explain reasons | Harder to summarise; interpretation may be subjective |
| Primary | Designed for the exact research aim | Can take time and effort to collect |
| Secondary | Quick, cheap and often large-scale | May not perfectly fit the new research question |
Numbers can still be weak evidence
Quantitative data can look “scientific”, but it may still be invalid if the question or measure is poor. For example, a rating scale about “happiness” may not really measure happiness well.
Descriptive statistics are simple ways of summarising a set of data. They describe what the data looks like; they do not prove that a result is significant.
A measure of central tendency is an average: it tells you where the “middle” or typical score is.
Mode
The mode is the most common score or category in a data set.
If data is grouped into intervals, the modal class is the interval with the highest frequency. For example, if the 11–15 score group has the most people in it, 11–15 is the modal class.
Median
The median is the middle value when all scores are placed in order from lowest to highest.
The median is useful when there is an extreme score because it is not pulled upwards or downwards as much as the mean.
Mean
The mean is found by adding all the scores together and dividing by the number of scores.
mean=total of scoresnumber of scores\text{mean} = \frac{\text{total of scores}}{\text{number of scores}}mean=number of scorestotal of scores
The mean uses all the data, which is a strength, but it can be distorted by an outlier: an extreme score that is very different from the rest.
Summarising a set of scores
A researcher records stress scores: 4, 5, 5, 7, 9, 18.
- Identify the mode by finding the most repeated score: 5 appears twice, so the mode is 5.
- Find the median by using the two middle scores, 5 and 7, because there is an even number of scores: 5+72=6\frac{5 + 7}{2} = 625+7=6.
- Find the mean by adding all scores and dividing by 6: 4+5+5+7+9+186=8\frac{4 + 5 + 5 + 7 + 9 + 18}{6} = 864+5+5+7+9+18=8.
- Compare the averages: 18 is an outlier, so the median of 6 may better represent the typical stress score than the mean of 8.
Range
The range is a measure of dispersion, or spread. It is calculated by subtracting the lowest score from the highest score.
For the stress scores above, the range is 18−4=1418 - 4 = 1418−4=14. A larger range means the scores are more spread out.
Mean, median or mode?
Use the mean when scores are fairly balanced, the median when there are outliers, and the mode when you need the most common category or response.
A ratio compares amounts. A fraction shows a part of a whole. A percentage shows a part of a whole out of 100.
percentage=partwhole×100\text{percentage} = \frac{\text{part}}{\text{whole}} \times 100percentage=wholepart×100
Converting between ratio, fraction and percentage
In a study, 18 participants improved after a memory strategy and 12 did not improve.
- Compare improved to not improved as a ratio: 18:12. Divide both sides by 6, giving 3:2.
- Write the improved group as a fraction of the whole sample: 1830\frac{18}{30}3018. Simplify by dividing top and bottom by 6, giving 35\frac{3}{5}53.
- Convert the fraction to a percentage: 1830×100=60%\frac{18}{30} \times 100 = 60\%3018×100=60%.
Percentages hide sample size
60% can mean 18 out of 30, but it could also mean 3 out of 5. Always check the actual number of participants if it is given.
A decimal is a number written using place value, such as 0.25. Standard form is a way of writing very large or very small numbers as a×10na \times 10^na×10n, where 1≤a<101 \le a < 101≤a<10.
A decimal place is a digit after the decimal point. Significant figures are the meaningful digits in a number, starting from the first non-zero digit.
Rounding and standard form
A questionnaire result is recorded as 0.004762.
- Move the decimal point so the first number is between 1 and 10: this gives 4.762×10−34.762 \times 10^{-3}4.762×10−3.
- Round to 2 significant figures by keeping 4 and 7, then using the next digit, 6, to round up: 4.8×10−34.8 \times 10^{-3}4.8×10−3.
- Round the original decimal to 3 decimal places by looking at the fourth decimal digit: 0.004762 becomes 0.005.
A normal distribution is a symmetrical, bell-shaped pattern of scores. Most scores cluster near the centre, with fewer scores at the extremes. In a perfect normal distribution, the mean, median and mode are in the same place.

Normal distributions are useful because they help researchers see what is typical and what is unusual. However, real psychology data may not be perfectly normal, especially with small or biased samples.
An estimate is an approximate value based on the data available. You may estimate from a graph, a table or a pattern in results.
A good estimate should be sensible, not wildly over-precise. Use words like “about” or “approximately” when the exact value cannot be read clearly.
Different displays suit different data. The aim is to make patterns clear without misleading the reader.

A frequency table, or tally chart, shows how often each score or category occurs. It is useful for organising raw counts before drawing a graph.
A bar chart shows separate bars for separate categories, such as different age groups or conditions. The bars should have gaps because the categories are separate.
A pie chart shows proportions of a whole. The slices should add up to 100%.
Calculating a pie chart sector
A class survey finds that 12 out of 40 students prefer visual revision methods.
- Work out the fraction of the whole group: 1240\frac{12}{40}4012.
- Convert this to a percentage: 1240×100=30%\frac{12}{40} \times 100 = 30\%4012×100=30%.
- Convert it to a pie chart angle: 1240×360=108\frac{12}{40} \times 360 = 1084012×360=108, so the sector is 108 degrees.
A histogram shows continuous grouped data, such as score intervals. The bars touch because the intervals join together.
A line graph shows change over a continuous variable, often time. It is good for trends, but joining points can imply there are meaningful values between the points.
A scatter diagram plots pairs of scores to show whether two variables are related. A positive correlation means both variables increase together. A negative correlation means one increases as the other decreases. Zero correlation means there is no clear relationship.
Bar chart or histogram?
Use a bar chart for separate categories and a histogram for continuous grouped data. The easiest visual clue is that histogram bars touch.
Correlation is not cause
A scatter diagram can show a relationship, but it cannot prove that one variable caused the other. A third variable may explain the pattern.
Reliability
Reliability means consistency. A reliable measure gives similar results when it is repeated in the same conditions.
- Internal reliability means consistency within a measure, such as whether questionnaire items all seem to measure the same thing.
- External reliability means consistency over time or across occasions, such as similar results if a test is repeated.
- Inter-rater reliability means different observers or raters record behaviour in a similar way.
Reliability is important because inconsistent findings are hard to trust. However, a study can be reliable but still invalid: it may consistently measure the wrong thing.
Validity
Validity means accuracy: whether a study or measure really measures what it claims to measure.
- Ecological validity means the setting, task or behaviour reflects real life.
- Population validity means the sample represents the target population, so findings can be generalised.
- Construct validity means the measure genuinely captures the psychological concept being studied, such as anxiety, memory or obedience.
Reliability versus validity
Reliability asks, “Is it consistent?” Validity asks, “Is it accurate?” You need both for strong psychological evidence.
Demand characteristics happen when participants guess the aim of a study and change their behaviour. The observer effect happens when people act differently because they know they are being watched. Social desirability happens when participants give answers that make them look good rather than fully honest.
Judging reliability and validity
A researcher asks students to rate how kind they are while their teacher watches. A second researcher observes the same students but records very different behaviours.
- The teacher watching may create social desirability, because students may give answers that make them seem kind.
- The measure may have low construct validity, because it may measure wanting approval rather than real kindness.
- The different observer records suggest low inter-rater reliability, so clearer behaviour categories would be needed.
Observer effect is not observer bias
The observer effect is about participants changing because they are watched. Observer bias is about the observer recording or interpreting behaviour unfairly.
Bias means a systematic unfairness or distortion in the research process. Bias can reduce validity and make conclusions less trustworthy.
- Gender bias: the sample, materials or interpretation favours one gender.
- Cultural bias: one culture’s norms are treated as if they apply to everyone.
- Age bias: findings from one age group are applied too widely.
- Experimenter bias: the researcher’s expectations influence how the study is run.
- Observer bias: the observer records behaviour in a way shaped by expectations.
- Bias in questioning: questions are leading, loaded or worded unevenly.
Spotting bias in a research design
A teacher studies phone use and sleep by asking only Year 10 students, “Don’t you agree that phones ruin sleep?”
- The sample may show age bias because only one year group is included.
- The wording shows bias in questioning because “Don’t you agree” and “ruin” push participants towards one answer.
- The teacher’s expectations could create experimenter bias, so an anonymous questionnaire with neutral wording would improve validity.
How to reduce bias
Use neutral questions, standardised instructions, anonymous responses, representative sampling and more than one observer where possible.
In the exam
- For calculation questions, show the method as well as the final answer, especially for means, percentages and pie chart angles.
- For graph questions, match the display to the data type: categories need bar charts; continuous grouped data needs histograms; relationships need scatter diagrams.
- For evaluation questions, use the key words precisely: reliability means consistency, validity means accuracy, and bias means systematic unfairness.
- When applying to a scenario, quote or refer to a detail from the source, then name the correct concept.
Check yourself
- What is the difference between quantitative and qualitative data?
- When might the median be a better average than the mean?
- How are demand characteristics different from social desirability?