x

Using genome projects (A-level only)

What you'll learn

  • What a genome project does and why many organisms’ genomes have been sequenced.
  • How genome data can be used to predict the proteome of simpler organisms.
  • Why the genome of a complex organism cannot be translated directly into “all its proteins”.
  • How genome projects can help identify antigens for vaccine production.

The starting point: DNA, genes and proteins

A cell’s DNA contains information in the order of its bases: adenine, thymine, cytosine and guanine. A gene is a length of DNA that codes for a polypeptide or a functional RNA molecule.

A protein’s amino acid sequence is ultimately determined by the sequence of bases in DNA. During gene expression, DNA is transcribed into messenger RNA, then translated at a ribosome into a polypeptide.

Definition

Genome and proteome

A genome is the complete set of DNA in an organism or cell. A proteome is the full range of proteins produced by a cell, tissue or organism at a particular time and under particular conditions.

This topic is part of gene expression because a genome is a set of instructions, but the proteome depends on which instructions are actually used.

What is a genome project?

A genome project is a scientific project that determines the base sequence of an organism’s genome. Sequencing projects have now read the genomes of a wide range of organisms, including humans, bacteria, viruses, plants and model organisms used in research.

Definition

DNA sequencing

DNA sequencing means determining the order of bases in a DNA molecule, usually recorded using the letters A, T, C and G.

Modern sequencing is highly automated. Machines generate very large amounts of sequence data, and computers are used to assemble, store, compare and analyse these data. This computer-based analysis of biological data is called bioinformatics.

Definition

Genome annotation

Genome annotation is the process of identifying meaningful features in a DNA sequence, such as genes, regulatory sequences and predicted protein-coding regions.

The overall idea is: sequence the genome, annotate it, then use the information to make biological predictions.

This overview shows why the link from genome to proteome is much more direct in simpler organisms than in complex eukaryotes.

Flow diagram showing how genome projects can predict proteins in simple organisms but not directly in complex eukaryotes

Using genomes of simpler organisms

In simpler organisms, especially many prokaryotes such as bacteria, a relatively large proportion of the genome codes for proteins. Their genes are often easier to identify because there is generally less non-coding DNA than in complex eukaryotes.

Scientists can search the genome for open reading frames, often shortened to ORFs.

Definition

Open reading frame

An open reading frame is a sequence of DNA that has the potential to code for a polypeptide, usually because it contains a start codon, a series of codons, and then a stop codon in the same reading frame.

Once a likely coding sequence has been identified, the genetic code can be used to predict the amino acid sequence of the protein. This helps build a predicted proteome for the organism.

Key Idea

Simple organisms

For simpler organisms, determining the genome can allow scientists to predict many of the protein sequences in the organism’s proteome.

Example

Predicting a short amino acid sequence

A short coding DNA sequence is: ATG GAA TTT TAA. A codon table gives: AUG = methionine, GAA = glutamic acid, UUU = phenylalanine, UAA = stop.

  1. Identify the triplets in the same reading frame: ATG, GAA, TTT, TAA.
  2. Convert the coding DNA sequence into the equivalent mRNA codons by replacing T with U: AUG, GAA, UUU, UAA.
  3. Use the codon table in order: AUG codes for methionine, GAA codes for glutamic acid, and UUU codes for phenylalanine.
  4. Stop translation at UAA, because it is a stop codon rather than an amino acid. The predicted polypeptide sequence is methionine–glutamic acid–phenylalanine.

Application: finding vaccine antigens

An antigen is a molecule that is recognised by the immune system and can stimulate an immune response. Many useful vaccine antigens are proteins found on the surface of a pathogen, because antibodies and immune cells can recognise them.

Genome projects can help vaccine development by allowing scientists to predict proteins made by a pathogen. They can then look for proteins that are likely to be exposed on the pathogen’s surface and suitable for use in a vaccine.

Definition

Antigen

An antigen is a molecule, often a protein or glycoprotein, that is recognised as foreign by the immune system and may trigger antibody production or other immune responses.

A useful vaccine antigen is often:

  • present on the pathogen’s surface
  • conserved across many strains of the pathogen
  • different from human proteins, reducing the risk of attacking human cells
  • important for the pathogen, so it is less likely to be lost by mutation
Example

Choosing a potential vaccine antigen

A bacterial genome project predicts three proteins. Protein A is a cytoplasmic enzyme found inside the bacterium. Protein B is a surface protein found in most strains. Protein C is a surface protein found in only one rare strain.

  1. Compare location first. Protein A is inside the bacterium, so it is less likely to be recognised easily by antibodies than a surface protein.
  2. Compare how widely the protein occurs. Protein B is found in most strains, while Protein C is found in only one rare strain.
  3. Choose the best candidate. Protein B is the strongest potential vaccine antigen because it is exposed on the surface and is shared by many strains.
Tip

Vaccine wording

If asked how genome projects help vaccine production, link the sequence data to predicted proteins, then to surface antigens, then to immune response.

Why complex organisms are harder

In more complex organisms, such as humans, knowing the genome sequence does not allow scientists to work out the full proteome easily.

There are several reasons.

Much DNA is non-coding

Non-coding DNA is DNA that does not code for the amino acid sequence of a polypeptide. This includes introns, some repeated sequences, and regulatory sequences.

This means a genome sequence contains far more than a simple list of protein-coding instructions.

Common Mistake

Non-coding does not mean useless

Do not describe all non-coding DNA as “junk”. Some non-coding DNA has important roles in controlling gene expression, chromosome structure or RNA production.

Genes can contain introns and exons

In eukaryotes, many genes contain exons, which are coding regions that remain in mature mRNA, and introns, which are removed during RNA splicing.

Different combinations of exons can sometimes be joined together. This is called alternative splicing, and it means one gene can help produce more than one protein.

Regulatory genes affect expression

A regulatory gene is a gene whose product controls the expression of other genes. For example, a regulatory gene may code for a transcription factor, which affects whether another gene is transcribed.

This matters because most cells in your body contain the same genome, but they do not all produce the same proteins. A muscle cell, neurone and lymphocyte all express different sets of genes, so they have different proteomes.

The proteome changes over time

The proteome is not fixed. It can change with:

  • cell type
  • stage of development
  • environmental conditions
  • hormones and signalling molecules
  • disease state
  • post-translational modification of proteins

Post-translational modification means changes made to a protein after translation, such as adding chemical groups or cutting the protein into an active form.

Key Idea

Complex organisms

In complex organisms, the genome shows the potential instructions, but the proteome depends on gene regulation, splicing, cell type and conditions.

Example

Explaining why genome data alone is not enough

A human liver cell and a human neurone contain essentially the same genome, but they have different proteomes.

  1. Start from the shared information. Both cells contain the same DNA sequence, so differences in protein production cannot usually be explained by different genomes.
  2. Apply gene regulation. Different regulatory proteins and regulatory sequences cause different genes to be transcribed in each cell type.
  3. Apply RNA processing. Some genes may be alternatively spliced, so different mature mRNA molecules can be produced from the same gene.
  4. Link to the proteome. Because different mRNAs are translated, and proteins may be modified after translation, the two cells produce different sets of proteins.

Sequencing methods keep changing

You do not need to memorise detailed sequencing methods for this specification point, but you should understand the trend: sequencing has become faster, cheaper, more automated and more dependent on computer analysis.

Early genome projects took many years and large international teams. Modern sequencing technologies can generate huge amounts of data much more quickly. However, producing a sequence is only the beginning. Scientists still need to identify genes, compare sequences, predict proteins and test predictions experimentally.

Common Mistake

Prediction is not proof

A genome sequence can suggest that a protein may exist, but experiments are often needed to confirm that the protein is actually produced, where it is found, and what it does.

Pulling the topic together

Genome projects are powerful because DNA sequence data can be stored, searched and compared. In simpler organisms, the path from DNA sequence to protein sequence is often fairly direct, so scientists can predict much of the proteome.

In complex organisms, the situation is more complicated. Non-coding DNA, regulatory genes, alternative splicing and different patterns of gene expression mean that the genome cannot simply be translated into one fixed proteome.

Exam technique

In the exam

  1. If asked about simpler organisms, explain that sequencing allows protein-coding genes to be identified and amino acid sequences to be predicted.
  2. If asked about vaccines, mention predicted pathogen proteins, especially surface antigens, and link these to triggering an immune response.
  3. If asked about complex organisms, do not just say “they have more DNA”; refer to non-coding DNA, regulatory genes and different gene expression.
Self review

Check yourself

  • Why is it easier to predict the proteome of many bacteria than the proteome of a human?
  • How can a genome project help scientists identify a possible vaccine antigen?
  • Why can two human cells with the same genome have different proteomes?

Recap questions

Test yourself with 5 quick questions on this guide. Answer them all correctly to complete it.

Previous

How was this guide?

Using genome projects (A-level only) Revision Guide

  1. A Level
  2. /Biology
  3. /Using genome projects (A-level only)