What you'll learn
- How to distinguish between a population, a census and a sample.
- How sampling methods can produce representative or biased data.
- How to use sample statistics to make informal inferences about a population.
- How sample size, variability and bias affect the reliability of a conclusion.
Populations and Samples
Statistical investigations usually begin with a question about a group of individuals or objects.
Population
A population is the complete set of people or items about which you want to obtain information. Each individual person or item is called a sampling unit.
For example, if a college wants to investigate how long its students spend travelling to college, the population is all students at that college. Each student is a sampling unit.
Sometimes it is possible to collect information from every member of the population.
Census
A census is an investigation in which data is collected from every member of the population.
A census can provide detailed information, but it may be expensive, slow or impractical. It is also impossible when testing destroys the item, such as measuring the lifetime of every light bulb manufactured by a company.
Instead, researchers usually investigate part of the population.
Sample
A sample is a subset of the population from which data is collected. The number of sampling units in the sample is called the sample size.
A sample is normally quicker and cheaper to investigate than the whole population. However, because it contains only part of the population, its results involve uncertainty.
From sample to population
You collect data from a sample, calculate useful statistics and then use those results to make an inference, meaning a reasoned conclusion, about the population.
Sampling Frames
To select a sample, you often need a list of the population.
Sampling frame
A sampling frame is a numbered or ordered list of all the sampling units in the population from which the sample is selected.
Examples include a school register, an electoral register or a database of customers.
A sampling frame should be accurate and up to date. If some members of the population are missing, duplicated or no longer eligible, the resulting sample may be biased.
Confusing the population with the sampling frame
The population is the entire group you want to investigate. The sampling frame is the list used to select members of that group. They are only equivalent when the list is complete and accurate.
Representative Samples and Bias
A good sample should reflect the important characteristics of the population.
Representative sample
A representative sample has characteristics that are reasonably similar to those of the population, particularly characteristics related to the investigation.
For example, a sample used to estimate support for extended library opening hours might need to represent students from different year groups. Sampling only final-year students could give a misleading result.
Bias
Bias is a systematic tendency for a sampling method or data-collection process to favour certain outcomes. A biased sample consistently over-represents or under-represents part of the population.
Bias is different from ordinary random variation. Random variation can make one well-selected sample slightly different from another. Bias pushes results in a particular direction.
Recognising sampling bias
A gym asks people leaving a Saturday morning fitness class whether the gym should offer more exercise classes.
- The target population is all gym members, but the sample contains only members who attended one particular exercise class.
- People already attending a class are likely to be more interested in exercise classes than members who mainly use the swimming pool or weights area.
- The sample is therefore unlikely to be representative, so it may overestimate support for additional classes.
Check who could be selected
When judging a sample, ask whether every relevant part of the population had a reasonable chance of being included.
Methods of Selecting a Sample
Simple Random Sampling
In a simple random sample, every possible sample of the required size has an equal chance of being selected. This can be achieved by assigning a number to each member of the sampling frame and using a random number generator.
This method reduces deliberate selection bias, although the selected sample may still differ from the population by chance.
Systematic Sampling
In systematic sampling, sampling units are selected at regular intervals from an ordered sampling frame. A random starting point should be chosen first.
If the population size is NNN and the required sample size is nnn, the approximate sampling interval is
k=Nn.k=\frac{N}{n}.k=nN.Selecting a systematic sample
A company has 1200 employees and wants a systematic sample of 80 employees.
-
Calculate the interval:
k=120080=15.k=\frac{1200}{80}=15.k=801200=15. -
Choose a random starting position from 1 to 15. Suppose position 8 is selected.
-
Select positions 8, 23, 38, 53 and so on, adding 15 each time until 80 employees have been chosen.
Patterns in an ordered list
Systematic sampling can be biased if the sampling frame has a repeating pattern that matches the sampling interval. For example, selecting every seventh entry from data arranged by day of the week could repeatedly select the same day.
Stratified Sampling
A population can often be divided into relevant groups called strata. For example, students might be divided by year group, and employees might be divided by department.
In a stratified sample, each stratum is represented in proportion to its size in the population. Members should then be selected randomly from within each stratum.
For a particular stratum,
sample from stratum=size of stratumpopulation size×total sample size.\text{sample from stratum} = \frac{\text{size of stratum}}{\text{population size}} \times \text{total sample size}.sample from stratum=population sizesize of stratum×total sample size.Calculating a stratified sample
A sixth form has 300 Year 12 students and 200 Year 13 students. A stratified sample of 50 students is required.
-
The total population is 300+200=500300+200=500300+200=500 students.
-
The number selected from Year 12 is
300500×50=30.\frac{300}{500}\times 50=30.500300×50=30. -
The number selected from Year 13 is
200500×50=20.\frac{200}{500}\times 50=20.500200×50=20. -
Therefore, randomly select 30 Year 12 students and 20 Year 13 students. This preserves the population ratio of 3 to 2.
Occasionally, a calculation does not produce a whole number. Sensible rounding is required, while ensuring that the rounded values add to the required total sample size.
Opportunity and Quota Sampling
In opportunity sampling, the researcher uses sampling units that are readily available. It is quick and convenient, but often biased because accessible individuals may not represent the population.
In quota sampling, the population is divided into groups and the researcher collects a fixed number from each group. Unlike stratified sampling, selection within each group is not necessarily random. This makes quota sampling easier without a complete sampling frame, but it allows interviewer choice to introduce bias.
Random selection matters
Stratified sampling controls how many people are selected from each group and uses random selection within those groups. Quota sampling controls the numbers but usually allows non-random selection.
Making Informal Inferences
A value calculated from a sample, such as a sample mean or sample proportion, is called a statistic. A numerical feature of the whole population is called a parameter.
You can use a sample statistic as an estimate of the corresponding population parameter. For example:
- the sample mean can estimate the population mean;
- the sample proportion can estimate the population proportion;
- the sample median can suggest a typical population value.
An informal inference is a conclusion based on the sample evidence without carrying out a formal statistical hypothesis test or constructing a confidence interval.
Estimating a population proportion
In a random sample of 150 students, 96 say that they use the college library at least once a week. The college has 1250 students.
-
Calculate the sample proportion:
p^=96150=0.64.\hat{p}=\frac{96}{150}=0.64.p^=15096=0.64. -
Use this as an estimate of the population proportion, so about 64% of all students may use the library at least once a week.
-
An informal estimate of the number of such students is
0.64×1250=800.0.64\times 1250=800.0.64×1250=800. -
Conclude that approximately 800 students may use the library weekly, while recognising that another random sample could produce a different estimate.
The symbol p^\hat{p}p^ is read as “p-hat” and denotes a sample proportion used to estimate a population proportion.
Reliability of an Inference
A conclusion from a sample is more convincing when:
- the sampling method minimises bias;
- the sample is representative of the population;
- the sample size is reasonably large;
- the data does not have extreme variability;
- the question and method of data collection are neutral.
Larger random samples generally show less sampling variability, meaning less variation between the results of different possible samples. However, a large biased sample can still produce a misleading conclusion.
Assuming large means reliable
Increasing the sample size reduces random sampling variation, but it does not remove systematic bias. A carefully selected sample of 200 may be more useful than a biased sample of 2000.
You should also avoid claiming certainty. A sample can suggest, indicate or provide evidence for a population conclusion, but it rarely proves that the conclusion is exactly true.
In the exam
- Identify the population, sampling units and sampling frame precisely rather than using vague phrases such as “everyone”.
- When evaluating a sample, explain which group is under-represented or over-represented and state how this could affect the result.
- For systematic sampling, calculate the interval and include a random starting point; for stratified sampling, show the proportional calculation for each stratum.
- Phrase an inference in context and acknowledge uncertainty: the sample evidence suggests a conclusion about the population, rather than proving it.
- Remember that a larger sample reduces sampling variability but does not correct a biased sampling method.
Check yourself
- What is the difference between a population, a sample and a sampling frame?
- How would you select a systematic sample of 40 people from a population of 600?
- Why can a large opportunity sample still give an unreliable inference?