What you'll learn
- How DNA sequencing has developed from Sanger sequencing to high-throughput sequencing.
- How PCR, electrophoresis and DNA profiling are used to analyse DNA.
- How genetic engineering moves genes between organisms using enzymes and vectors.
- Why gene manipulation has medical, agricultural and ethical importance.
The starting point: genomes, genes and information
A gene is a length of DNA that codes for a polypeptide or a functional RNA. A genome is all the DNA in an organism, including genes and non-coding regions.
Genome
A genome is the complete set of genetic material in an organism or cell.
DNA stores information in the sequence of its bases: adenine, thymine, cytosine and guanine. Changes in this sequence are mutations. Different versions of the same gene are alleles.
Your genotype is your genetic makeup, while your phenotype is your observable characteristics, produced by genotype and environmental effects.
DNA sequencing: reading the base order
DNA sequencing
DNA sequencing means finding the exact order of bases in a length of DNA.
Sequencing can identify genes, compare alleles and reveal evolutionary relationships. It also allows scientists to predict the amino acid sequence of proteins.
Sanger sequencing
Sanger sequencing is a chain-termination method. It uses:
- a single-stranded DNA template
- a primer, which is a short DNA sequence that binds to the template
- DNA polymerase, which builds a complementary strand
- normal DNA nucleotides
- fluorescently labelled ddNTPs, which are modified nucleotides that stop DNA synthesis because they lack the 3' OH group needed to form the next phosphodiester bond
This produces fragments of different lengths. Each fragment ends with a labelled base. The fragments are separated by capillary electrophoresis, and a detector reads the fluorescent labels in order.

High-throughput sequencing
High-throughput sequencing means sequencing very large numbers of DNA fragments in parallel. You do not need the detailed chemistry of these newer methods for OCR H420, but you should understand the impact: sequencing has become much faster, cheaper and more automated.
Why sequencing changed biology
Modern biology can now compare whole genomes, not just individual genes, so researchers can ask questions about disease risk, ancestry, pathogen spread and evolution at a genome-wide scale.
What sequence data can be used for
Genome-wide comparisons
Bioinformatics is the use of computer tools to store, search and analyse biological data. Computational biology uses mathematical and computational models to investigate biological questions.
Sequencing allows comparisons:
- between individuals, to find variants linked with disease risk or phenotype
- between species, to find evolutionary relationships
- between pathogen samples, to study epidemiology, which is the spread and distribution of disease in populations
A single nucleotide polymorphism, or SNP, is a one-base difference in DNA sequence between individuals. A genome-wide association study compares many SNPs across many genomes to look for associations between genotype and phenotype.
For evolutionary comparisons, scientists often compare homologous genes. Homologous genes are genes inherited from a common ancestor. More similar sequences usually suggest a more recent common ancestor, although many genes should be compared for reliable conclusions.
Predicting amino acid sequences
A codon is a sequence of three bases in mRNA that codes for one amino acid or a stop signal. The genetic code is the set of rules linking codons to amino acids.
If you know the DNA sequence of a gene, you can predict the mRNA sequence and then the amino acid sequence of the polypeptide.
Coding strand vs template strand
The coding strand has the same base sequence as mRNA except that DNA has T where mRNA has U. The template strand is complementary and antiparallel to the mRNA.
Predicting a polypeptide from a DNA sequence
A coding DNA strand is 5'-ATG CCA TTT TGA-3'. The relevant mRNA codons are: AUG = methionine, CCA = proline, UUU = phenylalanine, UGA = stop.
- Convert the coding DNA sequence into mRNA by replacing T with U: 5'-AUG CCA UUU UGA-3'.
- Split the mRNA into codons from the start of the coding sequence: AUG, CCA, UUU, UGA.
- Translate each codon until the stop codon is reached: methionine, proline, phenylalanine, then stop.
- The predicted polypeptide is methionine-proline-phenylalanine.
Synthetic biology
Synthetic biology
Synthetic biology is the design and construction of new biological parts, systems or organisms, often using DNA sequences designed on a computer and made artificially.
Sequencing made synthetic biology possible because scientists can read useful sequences, redesign them, synthesise DNA and insert it into cells. For example, microbes can be engineered to produce useful molecules, or genes can be redesigned to improve expression in a host cell.
PCR: amplifying DNA
PCR
The polymerase chain reaction, or PCR, is a technique used to make many copies of a specific DNA region.
PCR needs:
- template DNA
- primers that bind either side of the target region
- free DNA nucleotides
- a heat-stable DNA polymerase, such as Taq polymerase
- a thermocycler, which changes temperature automatically
Each cycle has three main stages:
- Denaturation at about 95 °C: hydrogen bonds between DNA strands break.
- Annealing at about 50–65 °C: primers bind to complementary sequences.
- Extension at about 72 °C: DNA polymerase builds new strands from the primers.
PCR and electrophoresis are often used together: PCR increases the amount of DNA, then electrophoresis separates fragments so they can be compared.

Calculating DNA copies after PCR
A PCR begins with 4 copies of a DNA fragment. Assume perfect doubling for 25 cycles.
- Use the ideal PCR model, where NNN is final copy number, N0N_0N0 is starting copy number and nnn is number of cycles: N=N02nN=N_0 2^nN=N02n.
- Substitute the values: N=4×225N=4 \times 2^{25}N=4×225.
- Calculate: 225=33 554 4322^{25}=33\,554\,432225=33554432, so N=4×33 554 432=134 217 728N=4 \times 33\,554\,432=134\,217\,728N=4×33554432=134217728 copies.
- State in standard form: about 1.34×1081.34 \times 10^81.34×108 copies.
PCR is not perfect forever
The formula N=N02nN=N_0 2^nN=N02n is an ideal model. In real PCR, amplification slows when primers, nucleotides or polymerase activity become limiting.
Electrophoresis: separating molecules
Electrophoresis
Electrophoresis separates charged molecules using an electric field.
DNA fragments are negatively charged because of their phosphate groups, so they move towards the positive electrode. In an agarose gel, smaller DNA fragments move further because they pass through the gel matrix more easily.
A DNA ladder contains fragments of known sizes, measured in base pairs, so unknown fragments can be estimated by comparison.
Electrophoresis can also separate proteins. Protein movement depends on size, shape and charge; in some methods, proteins are treated so separation is mainly by size.
Practical gel electrophoresis
Load samples into wells near the negative electrode, because DNA needs to travel towards the positive electrode. Use a loading dye so the sample sinks and you can track movement through the gel.
DNA profiling
DNA profiling
DNA profiling is the production of a DNA pattern that can be compared between individuals or samples.
DNA profiling often analyses short tandem repeats, or STRs. These are short, repeated DNA sequences. The number of repeats varies between individuals, so PCR and electrophoresis can produce a characteristic banding pattern.
Uses include:
- forensics, such as comparing crime-scene DNA with suspect DNA
- identification of remains or biological relationships
- disease-risk analysis, by detecting alleles or sequence variants associated with increased risk
Interpreting a DNA profile
A crime-scene sample has bands at 120 bp, 200 bp and 340 bp. Suspect A has bands at 120 bp, 200 bp and 340 bp. Suspect B has bands at 120 bp, 220 bp and 340 bp.
- Compare each band position with the DNA ladder to estimate fragment sizes.
- Check whether every crime-scene band is present in each suspect profile.
- Suspect A matches all three bands, while Suspect B has a different middle band.
- Conclude that Suspect A could be the source of the DNA, but the profile alone does not prove guilt because contamination, sample handling and probability must be considered.
Disease risk is not certainty
A DNA variant associated with disease may increase risk without guaranteeing disease. Many phenotypes are affected by multiple genes and environmental factors.
Genetic engineering: moving genes between organisms
Genetic engineering
Genetic engineering is the deliberate modification of an organism’s genome, often by inserting a gene from another organism.
A vector is a carrier used to transfer DNA into a host cell. A plasmid is a small circular DNA molecule found in bacteria, commonly used as a vector.
A restriction enzyme cuts DNA at a specific recognition sequence. Some cuts leave sticky ends, which are short single-stranded overhangs. DNA ligase joins DNA fragments by forming phosphodiester bonds. DNA made by joining DNA from different sources is recombinant DNA.

A typical method is:
- Isolate the desired gene.
- Cut the gene and plasmid with the same restriction enzyme to produce compatible ends.
- Use DNA ligase to insert the gene into the plasmid, forming a recombinant plasmid.
- Transfer the plasmid into host cells. Electroporation uses a brief electric pulse to create temporary pores in the cell membrane.
- Identify transformed cells and allow them to express the inserted gene.
Choosing a restriction enzyme
A student wants to insert a gene into a plasmid. Enzyme X cuts inside the coding sequence of the gene. Enzyme Y cuts on either side of the gene and also cuts the plasmid once.
- Reject Enzyme X because it would cut within the gene and could disrupt the amino acid sequence of the protein.
- Choose Enzyme Y because it can remove the intact gene and open the plasmid at a compatible site.
- Use DNA ligase after cutting, because complementary sticky ends can base-pair but ligase is needed to seal the sugar-phosphate backbone.
- Transform host cells with the recombinant plasmid, for example by electroporation.
Ethical issues in manipulating genomes
Genetic technologies can bring major benefits, but they also raise questions about safety, fairness, ownership and welfare.
Plants
Insect-resistant genetically modified soya may contain a gene for an insecticidal protein. Potential benefits include reduced crop loss, higher yields and reduced insecticide spraying. Concerns include effects on non-target insects, evolution of resistant pests, gene flow to wild relatives and dependence on seed companies.
Microorganisms and pathogens
Genetically modified microorganisms can produce useful products such as insulin or enzymes. Genetically modified pathogens can also be used in research to understand infection and develop vaccines or treatments. However, modified pathogens raise biosafety concerns, including accidental release and misuse.
Animals and humans
Pharming means genetically modifying animals to produce pharmaceutical proteins, for example in milk. This can make valuable medicines, but raises animal welfare questions and concerns about containment and purification.
In humans, genetic manipulation could treat disease, but also raises concerns about consent, inequality and possible non-medical enhancement.
Patenting and technology transfer
Patenting gives legal ownership over an invention, such as a genetic technology or engineered seed. Patents may encourage investment, but can restrict access.
Technology transfer means sharing knowledge, resources and practical ability so others can use a technology. For GM seed, a key ethical issue is whether poor farmers can access useful crops fairly, without being trapped by high costs or restrictions on saving seed.
Ethics needs balance
Strong ethical answers weigh benefits against risks and name who is affected: patients, farmers, consumers, ecosystems, researchers, companies and future generations.
Gene therapy
Gene therapy
Gene therapy is the treatment of disease by altering genetic material in a patient’s cells.
The basic principle is to introduce a functional allele, alter expression of a faulty gene, or modify cells so they perform a therapeutic role. Delivery usually needs a vector, such as a harmless virus or another DNA delivery system.
Somatic cell gene therapy
Somatic cells are body cells that are not involved in reproduction. In somatic cell gene therapy, only the treated tissues are altered. The change is not passed to offspring.
This is generally considered more ethically acceptable, but it may need repeated treatment if cells are replaced over time.
Germ line cell gene therapy
Germ line cells are cells that can pass genetic information to offspring, such as gametes or cells in an early embryo. Germ line gene therapy would affect all cells of the resulting individual and could be inherited by future generations.
This has greater potential to remove inherited disease from a family line, but it raises much stronger ethical concerns because future generations cannot consent and unintended effects could be inherited.
In the exam
- For method questions, write the sequence clearly: cut with restriction enzyme, join with DNA ligase, transfer using a vector, then identify transformed cells.
- For gel or DNA profile questions, compare every band, not just the most obvious one, and use cautious wording such as “could be the source”.
- For ethics questions, give both a benefit and a risk, and link each point to a named context such as GM soya, pharming, pathogens or gene therapy.
Check yourself
- Why do ddNTPs stop DNA synthesis during Sanger sequencing?
- Why do smaller DNA fragments travel further in an agarose gel?
- What is the key difference between somatic cell gene therapy and germ line cell gene therapy?