- How psychologists carry out experiments, interviews, questionnaires, observations, case studies and correlations.
- The strengths and weaknesses of each method, especially for reliability and validity.
- Which research method best suits different research objectives.
- How to apply methods to short GCSE scenarios and evaluate them clearly.
A research objective is the precise thing a psychologist wants to find out. The method should match the objective.
For example:
- “Does one thing cause another?” → often an experiment.
- “What do people think or feel?” → often an interview or questionnaire.
- “How do people behave naturally?” → often an observation.
- “What is happening in one rare or unusual case?” → often a case study.
- “Are two measured variables related?” → often a correlation.
Core research terms
- Reliability means consistency: would the method give similar results if repeated?
- Validity means accuracy: does the method measure what it claims to measure?
- Internal validity means confidence that the results were caused by the factor being studied, not by uncontrolled variables.
- Ecological validity means how far the findings reflect real-life behaviour.
- Quantitative data is numerical data, such as scores, ratings or frequencies.
- Qualitative data is descriptive data, such as spoken answers, written explanations or observations in words.
Ethics is not optional
When research involves people, the British Psychological Society’s Code of Ethics and Conduct expects researchers to consider informed consent, deception, protection from harm, right to withdraw, confidentiality and debriefing.
An experiment is a method where researchers compare conditions to see whether one factor affects another.
The independent variable, or IV, is the factor being changed or compared. The dependent variable, or DV, is the outcome being measured. In many experiments, researchers try to control extraneous variables, which are unwanted factors that could affect the DV.
Experiments are most suitable when the objective is to investigate cause and effect.
These three experiment types vary mainly in control, setting and whether the researcher manipulates the IV.

A laboratory experiment takes place in an artificial or highly controlled setting. The researcher manipulates the IV and measures the DV.
Strengths: high control improves internal validity, and standardised procedures can make the study reliable and easy to repeat.
Weaknesses: behaviour may be less natural, so ecological validity can be lower. Participants may also show demand characteristics, meaning they guess the aim and change their behaviour.
A field experiment takes place in a real-life setting, but the researcher still manipulates the IV.
Strengths: behaviour is often more natural, so ecological validity is usually higher than in a lab.
Weaknesses: there is less control over extraneous variables, so reliability and internal validity may be lower. Ethical issues can also be trickier if people do not realise they are being studied.
A natural experiment studies the effect of a naturally occurring IV. The researcher does not create or manipulate the IV.
Strengths: useful when manipulation would be unethical or impossible, such as studying the effects of a real-life event.
Weaknesses: researchers have less control and cannot randomly allocate people to conditions, so cause and effect is less certain.
Natural experiment versus naturalistic observation
A natural experiment compares a naturally occurring IV. A naturalistic observation simply watches behaviour in its normal setting. “Natural” does not mean the same thing in both labels.
Choosing an experiment type
A psychologist investigates whether a notification sound reduces memory-test scores. Volunteers complete a memory task in a quiet room. Half hear a notification sound and half do not.
- The objective is to test whether the notification sound causes a change in memory scores, so an experiment is suitable.
- The IV is whether the notification sound is present or absent, and the DV is the memory-test score.
- The researcher controls the room and task, so this is a laboratory experiment rather than a field experiment.
- A strength is high control, which improves internal validity; a weakness is that the quiet testing room may not reflect real phone use, lowering ecological validity.
An interview is a research method where participants answer questions verbally. Interviews are useful for studying thoughts, feelings and experiences.
A structured interview uses the same questions in the same order for every participant.
Strengths: this improves reliability because the procedure is standardised. Answers are easier to compare.
Weaknesses: participants may not be able to explain unexpected ideas in depth, so validity can be limited.
An unstructured interview is more like a guided conversation. The interviewer can ask follow-up questions.
Strengths: it can produce rich qualitative data and may improve validity because participants can explain their experiences fully.
Weaknesses: it is harder to repeat exactly, so reliability is lower. There is also more risk of interviewer bias, where the interviewer’s wording, tone or reactions influence answers.
Selecting an interview type
A researcher wants to understand how teenagers experience anxiety before exams.
- The objective is to explore personal experiences, not just count how many students feel anxious.
- An unstructured interview is suitable because follow-up questions can explore individual thoughts and feelings.
- This may increase validity because answers can be detailed and meaningful.
- The weakness is lower reliability because each interview may be slightly different.
A questionnaire, or survey, asks participants to answer written questions. It is a self-report method because people report their own behaviour, thoughts or feelings.
Questionnaires are useful for collecting data from large samples quickly.
An open question allows participants to answer in their own words.
Strengths: produces qualitative data and can reveal unexpected detail.
Weaknesses: answers can be hard to compare, reducing reliability.
A closed question gives fixed answers, such as “yes/no” or multiple choice.
Strengths: produces quantitative data that is easy to compare and summarise.
Weaknesses: participants may be forced into an answer that does not fully fit, reducing validity.
A rating scale asks participants to choose a point on a scale, such as 1 to 5 from “strongly disagree” to “strongly agree”.
Strengths: useful for measuring the strength of attitudes and producing numerical data.
Weaknesses: people may interpret the scale differently, or choose middle answers to avoid making a decision.
Questionnaires work best when anonymity matters
Anonymous questionnaires can reduce social desirability bias, where participants give answers that make them look better rather than answering honestly.
Choosing question formats
A survey investigates students’ sleep before exams.
- To find out how many hours students slept, a closed question is suitable because it produces clear quantitative data.
- To measure how worried students felt, a rating scale is suitable because it shows strength of feeling.
- To understand why students stayed up late, an open question is suitable because students can explain reasons in their own words.
- A mixed questionnaire can balance reliability from closed questions with validity from richer open answers.
An observation involves watching and recording behaviour. Researchers often use behavioural categories, which are clearly defined behaviours to count or describe.
Observations are suitable when the objective is to study what people actually do, rather than what they say they do.
- Naturalistic observation: behaviour is observed in its usual setting.
- Controlled observation: some conditions are arranged or standardised by the researcher.
- Overt observation: participants know they are being observed.
- Covert observation: participants do not know they are being observed.
- Participant observation: the researcher joins the group being studied.
- Non-participant observation: the researcher watches from outside the group.
Strengths: observations can have high validity because they record real behaviour. Naturalistic observations can have strong ecological validity.
Weaknesses: there may be observer bias, where the observer records behaviour in a subjective way. Covert observation can raise ethical issues because participants may not give informed consent.
Reliability can be improved using clear behavioural categories and more than one observer. If observers agree, this supports inter-observer reliability.
Classifying an observation
A researcher sits in a café and secretly counts how often customers say “thank you” to staff.
- The café is a real-life setting, so the observation is naturalistic.
- Customers do not know they are being observed, so it is covert.
- The researcher is not joining in as a customer in the interaction being studied, so it is non-participant.
- A strength is natural behaviour, improving ecological validity; a weakness is the ethical issue of no informed consent.
A case study is an in-depth investigation of one person, one small group or one unusual situation. A sample is the group of people studied; case studies usually use very small samples.
Case studies often use qualitative data, such as interviews, observations, diaries or detailed background information. They may also include test scores.
They are most suitable for rare, unusual or complex cases. In GCSE Psychology, rare memory cases such as Wilson, Kopelman and Kapur (2008) show why detailed case-study evidence can be valuable: it can reveal patterns that a large survey might miss.
Strengths: case studies provide rich detail and can have high validity because they look closely at real experiences.
Weaknesses: findings may not be generalisable, meaning they may not apply to other people. Reliability can also be low because the exact same case cannot usually be repeated.
When a case study fits
A psychologist studies one person with a rare memory disorder after brain damage.
- The case is unusual, so a large questionnaire may not capture the important detail.
- A case study allows interviews, memory tasks and observations to be combined.
- This can increase validity because the researcher builds a detailed picture of the person’s difficulties.
- The main weakness is limited generalisability because one person’s brain injury may not represent everyone with memory problems.
A correlation investigates whether two measured variables are related. In correlations, the variables are often called co-variables because neither is manipulated by the researcher.
Correlations use quantitative data because each participant needs paired numerical scores.
Scatter diagrams show the direction of the relationship visually.

A positive correlation means that as one co-variable increases, the other tends to increase.
A negative correlation means that as one co-variable increases, the other tends to decrease.
A zero correlation means there is no clear relationship between the two co-variables.
Strengths: correlations are useful when manipulation would be unethical or impossible. They can help make predictions.
Weaknesses: a correlation cannot prove cause and effect. A third variable may explain the relationship.
Correlation does not mean causation
If stress and poor sleep are correlated, you cannot automatically say stress causes poor sleep. Poor sleep might increase stress, or another factor such as illness could affect both.
Interpreting a correlation
A psychologist records students’ hours of gaming and sleep-quality scores. Students who game for longer tend to have lower sleep-quality scores.
- The two co-variables are gaming hours and sleep-quality score.
- As gaming hours increase, sleep-quality scores decrease, so the relationship is negative.
- The correct conclusion is that gaming time is associated with poorer sleep quality.
- It would be invalid to conclude that gaming causes poor sleep, because this is a correlation, not an experiment.
| Method | Best for objectives about... | Reliability and validity reminder |
|---|
| Laboratory experiment | Testing cause and effect under control | High reliability and internal validity, but lower ecological validity |
| Field experiment | Testing cause and effect in real settings | More realistic, but less control |
| Natural experiment | Studying naturally occurring differences | Useful when manipulation is unethical or impossible, but weaker control |
| Structured interview | Comparing verbal answers across people | More reliable, less depth |
| Unstructured interview | Exploring personal experiences | Rich validity, lower reliability |
| Questionnaire | Gathering self-report data from many people | Efficient and reliable if standardised, but may suffer from dishonest answers |
| Observation | Recording actual behaviour | Can be valid, but observer bias and ethics matter |
| Case study | Investigating rare or complex cases | Detailed and valid, but not very generalisable |
| Correlation | Checking whether two variables are related | Useful for prediction, but cannot show causation |
In the exam
- Name the method precisely: for example, say field experiment, covert naturalistic observation or structured interview, not just “study”.
- Link every strength or weakness to the scenario using reliability or validity: explain why control, realism, standardisation or bias matters.
- For correlations, always identify the two co-variables, state the direction, and avoid cause-and-effect language.
- If people are involved, add one relevant ethical issue such as consent, deception, harm, withdrawal, confidentiality or debriefing.
Check yourself
- What is the difference between a field experiment and a natural experiment?
- Why might an unstructured interview have high validity but low reliability?
- How can a correlation be useful even though it cannot prove cause and effect?