x

Statistical skills

What you'll learn

  • How to choose and calculate averages: mean, median, mode and modal class.
  • How to measure spread using range, quartiles, inter-quartile range and cumulative frequency.
  • How to calculate percentage change and interpret percentiles.
  • How to describe scatter graphs, draw lines of best fit, make predictions and spot misleading statistics.

Why statistical skills matter in Geography

Statistics help you turn messy real-world data into clear geographical evidence. You might use them for river fieldwork, urban quality-of-life surveys, climate data, development indicators, or case studies such as Lagos, Nepal 2015, or a UK city regeneration scheme. Exact figures vary by source, so the skill is not just remembering numbers — it is knowing what the numbers mean.

These skills are assessed across all three GCSE Geography papers, including your fieldwork and issue evaluation work.

Starting point: what is a data set?

A data set is a collection of values. For example, you might collect river depth measurements across a channel, pedestrian counts in a town centre, or rainfall totals for different months.

A variable is something that can be measured or recorded, such as rainfall in mm, population in millions, distance in km, or GNI per head in US$.

A frequency is how often a value or class occurs. If five days had rainfall between 10 and 19 mm, the frequency for that class is 5.

Definition

Grouped data and modal class

Grouped data is data placed into class intervals, such as 0–9 mm, 10–19 mm and 20–29 mm. The modal class is the class interval with the highest frequency. It tells you the most common group, not the exact most common value.

Measures of central tendency

A measure of central tendency is a single value that summarises the “typical” value in a data set. The three you need are mean, median and mode.

Mean

The mean is found by adding all the values and dividing by the number of values.

Mean=total of all valuesnumber of values\text{Mean}=\frac{\text{total of all values}}{\text{number of values}}Mean=number of valuestotal of all values​

The mean is useful when the data is fairly even, but it can be distorted by an outlier — a value that is unusually high or low.

Median

The median is the middle value when the data is placed in order from smallest to largest. It is useful when the data is skewed, such as house prices or GNI per head, because it is less affected by very large or very small values.

Mode

The mode is the value that occurs most often. For grouped data, use the modal class instead.

Example

Choosing an average for rainfall data

A student records daily rainfall in mm for seven days: 6, 4, 38, 7, 4, 2, 9.

  1. Put the values in order so the middle and repeated values are easier to see: 2, 4, 4, 6, 7, 9, 38.
  2. Calculate the mean: the total is 70 mm, so the mean is 70 divided by 7 = 10 mm.
  3. Find the median: there are seven values, so the 4th value is the middle value. The median is 6 mm.
  4. Find the mode: 4 mm appears twice, more than any other value, so the mode is 4 mm.
  5. Compare the averages: the mean is higher than the median because the 38 mm day is an outlier. The median may better represent a “typical” day.
Tip

Choosing the best average

Use the mean when values are fairly evenly spread, the median when there are outliers, and the mode/modal class when you want the most common value or category.

Measures of spread

A measure of spread shows how varied the data is. In Geography, this helps you compare places or time periods. For example, two rivers may have the same mean depth, but one may vary much more across its channel.

Range

The range is the difference between the highest and lowest values.

Range=highest value−lowest value\text{Range}=\text{highest value}-\text{lowest value}Range=highest value−lowest value

A large range means the data is more spread out. However, the range can be strongly affected by one extreme outlier.

Quartiles and inter-quartile range

Quartiles divide ordered data into four equal parts.

  • The lower quartile, written as Q1Q_1Q1​, is about the 25th percentile.
  • The median, or Q2Q_2Q2​, is the 50th percentile.
  • The upper quartile, written as Q3Q_3Q3​, is about the 75th percentile.

The inter-quartile range shows the spread of the middle 50% of the data.

IQR=Q3−Q1IQR=Q_3-Q_1IQR=Q3​−Q1​

The inter-quartile range is useful because it ignores the most extreme quarter at each end.

Cumulative frequency

Cumulative frequency means a running total of frequencies. It is often used to estimate the median, quartiles and percentiles from grouped data.

On a cumulative frequency graph:

  • Q1Q_1Q1​ is found at one quarter of the total frequency.
  • The median is found at half of the total frequency.
  • Q3Q_3Q3​ is found at three quarters of the total frequency.

Annotated cumulative frequency graph showing Q1, median, Q3 and inter-quartile range

Example

Reading spread from a cumulative frequency graph

The cumulative frequency graph shows rainfall for 40 days.

  1. Work out the positions: Q1Q_1Q1​ is at 10 days, the median is at 20 days, and Q3Q_3Q3​ is at 30 days.
  2. Read across from cumulative frequency 10 to the curve, then down to the rainfall axis. This gives Q1≈20Q_1 \approx 20Q1​≈20 mm.
  3. Repeat for cumulative frequency 20 and 30. The median is about 35 mm and Q3≈49Q_3 \approx 49Q3​≈49 mm.
  4. Calculate the inter-quartile range: IQR=49−20=29IQR=49-20=29IQR=49−20=29 mm.
  5. Interpret it: the middle 50% of days had rainfall spread across about 29 mm.
Common Mistake

Quartile methods can vary slightly

For a short list of raw data, different textbooks sometimes handle the median value slightly differently when finding quartiles. In GCSE graph questions, use the graph method: total frequency, then one quarter, one half and three quarters.

Percentage increase and decrease

A percentage change shows how much a value has changed relative to its original value. This is useful for population growth, tourism changes, deforestation rates, energy use, or changes in flood risk.

Percentage change=new value−old valueold value×100\text{Percentage change}=\frac{\text{new value}-\text{old value}}{\text{old value}}\times 100Percentage change=old valuenew value−old value​×100

If the answer is positive, it is a percentage increase. If it is negative, it is a percentage decrease.

Example

Calculating percentage increase

A settlement’s population rises from 80,000 to 100,000.

  1. Find the change: 100,000 minus 80,000 = 20,000.
  2. Compare the change with the original value: 20,000 divided by 80,000 = 0.25.
  3. Convert to a percentage: 0.25 multiplied by 100 = 25%.
  4. Interpret the result: the population increased by 25%.
Common Mistake

Dividing by the wrong number

For percentage change, always divide by the original value, not the new value. The question is asking how large the change is compared with the starting point.

Percentiles

A percentile tells you the value below which a certain percentage of data lies.

For example, if a river discharge is at the 90th percentile, it means about 90% of recorded discharges are lower than or equal to that value, and about 10% are higher.

Percentiles are useful when comparing a place with a wider data set. For example, a country in the 80th percentile for GNI per head is wealthier than about 80% of countries in that data set. The exact ranking depends on the source and year.

Key Idea

Percentiles are about position

A percentile is not the same as a percentage change. It tells you where a value sits within a ranked data set.

Bivariate data and scatter plots

Bivariate data means data with two variables measured for the same cases. For example:

  • distance from the city centre and average house price
  • GDP per head and life expectancy
  • river velocity and distance downstream
  • beach sediment size and distance along a transect

A scatter plot shows bivariate data as points on a graph. Each point represents one place, time, site or country.

A relationship is the pattern between the two variables.

  • A positive relationship means both variables tend to increase together.
  • A negative relationship means one variable tends to decrease as the other increases.
  • No clear relationship means the points show no obvious pattern.

Scatter plot with line of best fit, interpolation and extrapolation

Lines of best fit and trend lines

A line of best fit is an estimated line drawn through the middle of the points on a scatter plot. It should have roughly similar numbers of points on each side. It does not need to pass through the origin or through every point.

A trend line shows the general direction of the data. It may be straight or curved, depending on the pattern.

Example

Predicting from a line of best fit

A scatter plot shows distance from a city centre and average house price.

  1. Identify the direction of the relationship: the line slopes downwards, so this is a negative relationship.
  2. For a value inside the data range, such as 6 km, read up to the line of best fit and across to the y-axis. The estimated house price is about £265,000.
  3. This is interpolation because 6 km lies within the range of the plotted data.
  4. If the line is extended to 20 km, the prediction is about £50,000, but this is extrapolation because it is beyond the observed data.
  5. Treat the extrapolated prediction as less reliable because the pattern may not continue outside the data range.
Common Mistake

Correlation is not causation

A scatter graph can show that two variables are related, but it does not prove that one directly causes the other. For example, house prices may be affected by distance from the centre, transport links, schools, employment and environmental quality.

Weaknesses in selective statistical presentation

Selective statistical presentation means data is chosen or shown in a way that supports a particular viewpoint, while hiding or weakening other evidence.

You may be asked to identify weaknesses in a graph, map, table, report extract or infographic.

Look out for:

  • a very small sample size
  • data from only one time, season or location
  • percentages given without raw numbers
  • totals used instead of rates per person or per area
  • mean values hiding inequality or outliers
  • a graph with a truncated axis that exaggerates change
  • choropleth classes chosen to make differences look bigger or smaller
  • missing source, date or method
  • comparing places with very different populations or sizes
Example

Spotting misleading totals

A map claims Local Authority A has a bigger crime problem than Local Authority B because A has 2,000 recorded incidents and B has 800.

  1. Check whether totals are fair to compare: A has a population of 200,000, while B has a population of 40,000, so the populations are very different.
  2. Calculate rates per 1,000 people: A has 10 incidents per 1,000 people, while B has 20 incidents per 1,000 people.
  3. Reinterpret the data: the total number is higher in A, but the rate is higher in B.
  4. Identify the weakness: the presentation is selective because it uses totals rather than comparable rates.
Tip

Challenge the statistic

Ask: Compared with what? Over what time period? From what source? For which area? Using totals or rates? These questions help you evaluate statistical evidence quickly.

Exam technique

In the exam

  1. Choose a statistic that matches the question: use averages for typical values, spread for variation, percentage change for change over time, and scatter graphs for relationships.
  2. Always include units where relevant, such as mm, km, £000s, millions or US$ per head.
  3. When describing a relationship, give the direction, strength, evidence from the graph, and any outliers.
  4. For percentage change, divide by the original value and state whether it is an increase or decrease.
  5. When evaluating data presentation, comment on sample size, scale, source, missing data and whether totals or rates are being used.
Self review

Check yourself

  • When would the median be a better average than the mean in a geographical data set?
  • How do you find the inter-quartile range from a cumulative frequency graph?
  • Why is extrapolation from a scatter graph less reliable than interpolation?

Recap questions

Test yourself with 5 quick questions on this guide. Answer them all correctly to complete it.

PreviousNext

How was this guide?

Statistical skills Revision Guide

  1. GCSE
  2. /Geography
  3. /Statistical skills