Skip to content
MathsGenie logo
Open app

Course home

  1. A Level
  2. Biology AQA
  3. Revision guides

Investigating diversity

Welcome to the final part of the genetic diversity topic! You already know what genetic diversity is—the number of different alleles of genes in a population. But how do biologists actually measure it in the real world?

What you'll learn:

  • Why we have moved from observing physical traits to directly sequencing DNA and proteins.
  • How to interpret molecular sequences to figure out how closely related two species are.
  • How to collect valid data using random sampling.
  • How to interpret mean values and standard deviations to draw valid conclusions about variation.

The Shift in Investigating Diversity

Historically, if scientists wanted to figure out how much genetic diversity existed in a population, or how closely related two species were, they had to rely on what they could see.

The old way: Observable characteristics

The traditional method was to infer DNA differences by looking at observable characteristics (the organism's phenotype). If two organisms looked very similar, scientists assumed their DNA must be very similar.

However, there are two major problems with inferring genetic diversity from observable traits alone:

Definition

Polygenic inheritance

Many observable characteristics are polygenic, meaning they are coded for by more than one gene. This creates continuous variation, making it very hard to tease apart discrete genetic differences just by looking.

  1. Polygenic traits: Because many genes influence a single trait (like height or coat colour), it is impossible to trace a physical difference back to a single specific mutation in the DNA.
  2. Environmental influences: An organism's phenotype is a combination of its genotype and its environment. For example, a plant might be genetically capable of growing tall, but if it grows in poor, nutrient-deficient soil, it will be short. The environment masks the true genetic potential.

The new way: Gene technology

Gene technology has completely changed how we investigate diversity. We no longer have to guess what the DNA looks like based on a plant's height or an animal's fur. We can extract the genetic material and read the sequence directly.


Interpreting Molecular Data

Today, genetic diversity within a species, or evolutionary relationships between different species, can be measured by comparing four main things:

  1. The frequency of measurable characteristics (the modern, quantitative version of the old way).
  2. The base sequence of DNA.
  3. The base sequence of mRNA.
  4. The amino acid sequence of the proteins encoded by DNA and mRNA.

Comparing base and amino acid sequences

When one species gives rise to another during evolution, the genome of the new species will initially be very similar to the original. Over time, mutations occur. The more time that passes since two species shared a common ancestor, the more mutations will build up in their DNA.

Key Idea

Evolutionary relationships

Closely related species have very similar DNA base sequences. Distantly related species have accumulated more mutations over time, resulting in more differences in their DNA, mRNA, and amino acid sequences.

  • DNA base sequence: This is the most accurate method. By reading the exact order of A, T, C, and G, scientists can count the exact number of base differences between two organisms.
  • mRNA base sequence: Because mRNA is transcribed directly from DNA, comparing mRNA sequences also provides an excellent measure of diversity.
  • Amino acid sequence: You can compare the sequence of amino acids in a specific protein (like haemoglobin or cytochrome c) that is shared between the organisms.

Comparing sequences

Tip

The degenerate code

Remember that the genetic code is degenerate (multiple triplet codes can code for the same amino acid). This means a mutation can change the DNA base sequence without changing the amino acid sequence. Therefore, comparing DNA base sequences is always more sensitive and accurate than comparing amino acid sequences.


Quantitative Investigations of Variation

When investigating variation within a single species, biologists usually measure quantitative (numerical) characteristics, like the length of leaves on an oak tree or the mass of mice in a field.

It is practically impossible to measure every single individual in a population. Instead, we have to take a sample. To ensure your sample data accurately represents the whole population, it must be collected randomly.

Taking a random sample

If you choose which leaves to measure yourself, you might naturally pick the largest ones, or the ones easiest to reach. This introduces bias. To eliminate human bias, we use a random sampling method.

A standard method for random sampling in a field involves:

  1. Laying out two long tape measures at right angles to create a grid over the area.
  2. Using a random number generator (on a calculator or computer) to generate two coordinates.
  3. Placing a quadrat exactly at those coordinates and collecting the data.

Analyzing Quantitative Data

Once you have collected your random sample, you will have a large set of numbers. To make sense of them, you need to calculate two vital statistics.

The Mean

The mean is the average value. You calculate it by adding up all your measured values and dividing by the total number of measurements taken. The mean is useful for comparing the average sizes of two different populations, but it doesn't tell you anything about how spread out the data is.

Standard Deviation

Standard deviation tells you about the spread of values around the mean.

  • A small standard deviation means most of the data points are clustered closely around the mean. The data is highly reliable.
  • A large standard deviation means the data points are spread widely around the mean. There is a lot of variation in the sample.
Common Mistake

Calculating Standard Deviation

You will not be asked to calculate the standard deviation from scratch in your written A-Level exams. However, you will be expected to know what it means and how to interpret it when given the value.

Interpreting overlap

In the exam, you will frequently be given a bar chart or a table showing the means of two groups along with their standard deviations. You need to check if the standard deviations of the two groups overlap.

Standard Deviation Overlap

Common Mistake

Confusing chance and significance

A very common mistake is stating "the results are significant" without explaining why. Always link your conclusion back to the standard deviation overlap and the concept of chance.

Example

Interpreting standard deviation data

A student investigated the variation in leaf length between two different populations of plants.

  • Population A had a mean leaf length of 45.245.245.2 mm with a standard deviation of 3.13.13.1 mm.
  • Population B had a mean leaf length of 52.052.052.0 mm with a standard deviation of 2.82.82.8 mm. Based on this data, is the difference in mean leaf length likely to be significant?
  1. Calculate the spread (Mean±SD\text{Mean} \pm \text{SD}Mean±SD) for Population A.
Lower limit=45.2−3.1=42.1 mmUpper limit=45.2+3.1=48.3 mm\begin{aligned} \text{Lower limit} &= 45.2 - 3.1 = 42.1 \text{ mm} \\ \text{Upper limit} &= 45.2 + 3.1 = 48.3 \text{ mm} \end{aligned}Lower limitUpper limit​=45.2−3.1=42.1 mm=45.2+3.1=48.3 mm​

The standard deviation range for Population A is 42.1 mm42.1 \text{ mm}42.1 mm to 48.3 mm48.3 \text{ mm}48.3 mm.

  1. Calculate the spread (Mean±SD\text{Mean} \pm \text{SD}Mean±SD) for Population B.
Lower limit=52.0−2.8=49.2 mmUpper limit=52.0+2.8=54.8 mm\begin{aligned} \text{Lower limit} &= 52.0 - 2.8 = 49.2 \text{ mm} \\ \text{Upper limit} &= 52.0 + 2.8 = 54.8 \text{ mm} \end{aligned}Lower limitUpper limit​=52.0−2.8=49.2 mm=52.0+2.8=54.8 mm​

The standard deviation range for Population B is 49.2 mm49.2 \text{ mm}49.2 mm to 54.8 mm54.8 \text{ mm}54.8 mm.

  1. Compare the ranges to check for overlap. The highest value in Population A's spread (48.3 mm48.3 \text{ mm}48.3 mm) is lower than the lowest value in Population B's spread (49.2 mm49.2 \text{ mm}49.2 mm). The ranges do not cross over.

  2. State your conclusion. Because the standard deviations do not overlap, the difference in the mean leaf lengths is likely to be significant (it is unlikely to be due to chance).


Exam technique

In the exam

When answering questions about molecular phylogeny or statistical interpretation, keep these key phrases ready:

  1. Comparing sequences: State that "more similar DNA/amino acid sequences indicate a more recent common ancestor".
  2. Avoiding bias: If asked how to ensure a sample is valid, write "use a random number generator to select coordinates" rather than just saying "choose randomly".
  3. Statistical conclusions: If SDs overlap, write: "The standard deviations overlap, so the difference between the means is likely due to chance and is not significant."
  4. If SDs do NOT overlap, write: "The standard deviations do not overlap, so the difference between the means is likely significant and not due to chance."
Self review

Check yourself

  • Why is comparing DNA base sequences a more accurate measure of genetic diversity than comparing observable characteristics?
  • Why might two species have identical amino acid sequences for a specific protein, but slightly different DNA base sequences for the gene encoding it?
  • Describe exactly how you would position a quadrat randomly in a field to avoid bias.
  • If Population X has a mean of 151515 and an SD of 444, and Population Y has a mean of 202020 and an SD of 555, do their standard deviations overlap? What does this tell you about the difference in their means?
Recap questions

1 of 5

Two plants are grown in very different soils and end up with different heights. Which comparison would give the clearest evidence about how genetically similar they really are?

PreviousNext

How was this guide?

Teach Genie

Review Investigating diversity by teaching Genie

Teach it back in your own words, spot gaps, and remember it better.

Start teaching
Genie and Baby Genie

Lesson

Recap your knowledge with an interactive lesson

7 minute activity

Start lesson

Investigating diversity means measuring how much genetic variation exists within a population, or how closely related two species are. Older studies often relied on observable characteristics because phenotype is easier to record than DNA.

That approach is limited because many characteristics are polygenic, so one visible trait may be controlled by several genes. It is hard to match a difference in height, mass, or leaf shape to one specific DNA change.

Phenotype is also influenced by the environment, so similar-looking organisms do not always have similar genotypes. Modern gene technology lets biologists compare DNA, mRNA, and proteins directly, which gives a much more accurate picture.

Flashcards

Remember key concepts with flashcards

22 flashcards

Practice flashcards

What did scientists traditionally use to infer DNA differences?

Investigating diversity Revision Guide

  1. A Level
  2. /Biology
  3. /Investigating diversity