What you'll learn
- How psychologists turn behaviour into behavioural categories that can be recorded.
- The difference between event sampling and time sampling.
- How to apply these ideas to observation scenarios.
- How to evaluate observational design using reliability, validity and ethics.
Why observational design matters
An observation is a research method where psychologists record behaviour as it happens, rather than asking people what they think or remember.
Observational design
Observational design is the plan for how behaviour will be defined, recorded and sampled during an observation. In AQA terms, the key design decisions are behavioural categories, event sampling and time sampling.
A good observational design makes behaviour more scientific: it reduces guesswork, helps different observers record behaviour consistently, and makes the study easier to replicate.
The flowchart below shows the usual sequence: decide what you want to study, define the behaviour clearly, choose categories, then decide whether to record every event or record at set times.

Operationalising behaviour
Before you can observe behaviour, you must make vague ideas measurable.
Operationalisation
Operationalisation means defining a concept in clear, observable and measurable terms, so researchers know exactly what counts as that behaviour.
For example, “aggression” is too vague on its own. One observer might count sarcasm, another might only count physical violence. A better operational definition would be: “hitting, kicking, pushing, grabbing an object from another child, or shouting an insult.”
This matters because observations can easily become subjective. The more precisely a behaviour is operationalised, the more likely observers are to record the same thing.
Behavioural categories
Behavioural categories
Behavioural categories are the specific behaviours that observers look for and record during an observation.
They are usually written on a coding scheme or recording schedule: a prepared sheet where observers tick, tally or note behaviours when they occur.
Good behavioural categories should be:
- Objective: based on visible or audible behaviour, not the observer’s opinion.
- Observable: something the observer can actually see or hear.
- Mutually exclusive: one behaviour should not fit into two categories at the same time.
- Exhaustive: the categories should cover all relevant behaviours, or include an “other” category.
- Clearly worded: another observer should understand them in the same way.
Good behavioural categories
A strong behavioural category describes what the participant does, not what the observer thinks the participant feels or intends.
For example, “being rude” is weak because it requires interpretation. “Interrupts another person while they are speaking” is stronger because it is directly observable.
Improving behavioural categories
A researcher wants to observe children’s playground behaviour and begins with the categories “aggressive”, “friendly” and “naughty”.
-
Identify the problem with the original categories. “Aggressive”, “friendly” and “naughty” are vague labels. They depend on the observer’s judgement, so two observers may code the same action differently.
-
Replace labels with observable behaviours. “Aggressive” could become “hits”, “pushes”, “kicks” and “shouts an insult”. “Friendly” could become “shares equipment”, “invites another child to play” and “helps another child”.
-
Check that categories do not overlap. “Shouting” could be coded as both aggressive and disruptive, so the category should be more specific, such as “shouts an insult at another child”.
-
Check whether important behaviours are missing. If the researcher also wants to record neutral behaviour, they might add “plays alone” or “no interaction” so the coding scheme is more exhaustive.
Using mental states as categories
Avoid categories like “anxious”, “jealous” or “attention-seeking” unless they are operationalised into observable actions. Observers can record behaviour, but they cannot directly see a participant’s private thoughts.
Sampling behaviour
Once the behavioural categories are ready, the researcher must decide when to record behaviour. In AQA, the two key sampling methods are event sampling and time sampling.
Event sampling
Event sampling
Event sampling is when the observer records every time a target behaviour occurs during the observation period.
This is useful when the behaviour is a discrete event, meaning it has a clear beginning and end. Examples include:
- a child hits another child;
- a student raises their hand;
- a driver uses their phone;
- a patient presses a help button.
Event sampling is especially suitable when the researcher wants to know how often a behaviour happens.
Strengths of event sampling
Event sampling can produce clear quantitative data, such as the number of times a behaviour occurred. This makes it useful for comparisons between individuals, groups or conditions.
It is also good for behaviours that are important but not constant. For example, if a researcher is studying aggressive incidents, it makes sense to record each incident rather than only checking occasionally.
Limitations of event sampling
Event sampling can be difficult if behaviours happen very quickly, happen at the same time, or have unclear boundaries. An observer may miss events if too much is happening at once.
It also records frequency better than duration or context. For example, knowing that a child pushed another child three times does not tell you how long the conflict lasted or what caused it.
Time sampling
Time sampling
Time sampling is when the observer records behaviour at set time intervals, such as every 30 seconds or every 2 minutes.
Time sampling is useful when behaviour is ongoing rather than a single clear event. Examples include:
- whether a student is on-task or off-task during a lesson;
- whether a child is playing alone, in parallel, or with others;
- whether a patient is awake, asleep or interacting with staff.
In instantaneous time sampling, the observer records what is happening at the exact moment the interval occurs. In broader interval sampling, the observer may record whether a behaviour occurred at any point during that interval. At A-Level, the key idea is that behaviour is recorded according to time periods rather than every occurrence.
Strengths of time sampling
Time sampling makes observations more manageable, especially when there are many participants or behaviours. The observer does not have to record continuously, so they may be less overwhelmed.
It can also give a useful snapshot of how behaviour is distributed across a period of time, such as how much of a lesson a student appears to spend on-task.
Limitations of time sampling
Time sampling can miss short behaviours that occur between intervals. If a child hits another child just after one observation point and stops before the next, it may not be recorded at all.
The choice of interval also matters. Very long intervals may miss too much detail, while very short intervals can become difficult to manage.
Interval length matters
Time sampling is only as good as the interval chosen. If the behaviour is brief and rare, long intervals are likely to produce an unrepresentative record.
Choosing a sampling method
A psychologist is deciding how to observe behaviour in a classroom.
-
Match the method to the type of behaviour. If the psychologist wants to count how often students call out without raising their hand, this is a discrete event with a clear start and end.
-
Match the method to the research aim. If the aim is to measure frequency, event sampling is better because every call-out can be tallied.
-
Consider a different aim. If the psychologist wants to estimate how much of the lesson students spend on-task, time sampling is more suitable because “being on-task” is an ongoing state.
-
Justify the decision. Event sampling fits “how many times did it happen?”, while time sampling fits “what is happening at regular points across the observation period?”
Event sampling vs time sampling
| Feature | Event sampling | Time sampling |
|---|---|---|
| What is recorded? | Every occurrence of a target behaviour | Behaviour at set time intervals |
| Best for | Clear, discrete behaviours | Ongoing states or frequent behaviour |
| Main data produced | Frequency counts | Snapshots or estimates over time |
| Main weakness | Can miss context or become overwhelming | Can miss behaviours between intervals |
Quick distinction
Event sampling follows the behaviour. Time sampling follows the clock.
Reliability and validity in observational design
Reliability
Reliability means consistency. In observations, this often means different observers would record the same behaviour in the same way.
A major way to improve reliability is to use clear behavioural categories. If two observers are watching the same lesson, they should both know exactly what counts as “calling out” or “helping another student”.
Validity
Validity means whether the observation is really measuring what it claims to measure.
A coding scheme can be reliable but still not valid. For example, if “anxiety” is only measured by “biting nails”, observers may record it consistently, but the measure may miss many anxious behaviours and include people who bite nails for other reasons.
A pilot study is a small trial run of the procedure before the main study. It can help researchers test whether the categories are clear, whether the intervals are manageable, and whether observers need more training.
AO3: Evaluating observational design
Strengths
Well-designed observations can produce detailed evidence of real behaviour. This is useful because people may not accurately report what they do, especially if the behaviour is embarrassing, automatic or socially undesirable.
Behavioural categories make observations more objective. They also allow behaviour to be turned into numerical data, such as frequencies or proportions, which can be compared across groups.
Clear sampling methods also improve replication. Another researcher can repeat the observation using the same categories and intervals.
Weaknesses
Observations can still suffer from observer bias, where the observer’s expectations influence what they notice or record. For example, if an observer expects boys to be more aggressive, they may unconsciously notice boys’ rough play more than girls’ rough play.
Observations may also be reductionist. Complex behaviour is broken down into categories, which can remove meaning and context. A tally of “pushes” does not explain whether the push was hostile, playful, accidental or defensive.
Sampling methods also create limits. Event sampling may miss duration and context, while time sampling may miss brief but important behaviours.
Ethical considerations
Observations must still follow ethical principles, especially when people are being watched without full awareness.
Key ethical issues include:
- Consent: participants should usually agree to take part, especially in controlled or private settings.
- Deception: covert observation may involve deception because participants do not know they are being studied.
- Right to withdraw: this is harder in covert naturalistic observations because participants may not know they are involved.
- Protection from harm: researchers should avoid causing distress or embarrassment.
- Confidentiality: names and identifying details should be removed from records.
- Debrief: where possible, participants should be told the aim of the study afterwards.
Observing behaviour in a public place may reduce some consent problems, but researchers should still avoid recording private, sensitive or identifiable information without justification.
In the exam
-
Use the exact terms: behavioural categories, event sampling and time sampling. Define them clearly before evaluating.
-
Apply to the scenario: if the question describes a behaviour, explain why it is better suited to event sampling or time sampling.
-
Evaluate specifically: do not just say “it is reliable”. Explain how clear categories, mutual exclusivity or fixed intervals improve reliability or validity.
-
Mention ethics where relevant: especially consent, confidentiality and deception in covert or naturalistic observations.
Check yourself
- What makes a behavioural category objective, mutually exclusive and exhaustive?
- Why would event sampling be better than time sampling for recording aggressive incidents?
- What is one limitation of time sampling when behaviours are brief?
