What you'll learn
- What “reinforcement” means in operant conditioning, and how it differs from punishment.
- The difference between primary and secondary reinforcement.
- How schedules of reinforcement affect response rate and resistance to extinction.
- How to apply reinforcement properties to examples such as token economies, education, gambling, and animal learning.
The starting point: operant conditioning
Operant conditioning is learning through consequences. A behaviour is more or less likely to happen again depending on what follows it.
For example, if a rat presses a lever and receives food, it may press the lever more often in future. The behaviour “operates” on the environment, and the consequence changes future behaviour.
Reinforcement
Reinforcement is any consequence that increases the likelihood of a behaviour being repeated.
This definition is important: a consequence only counts as reinforcement if it actually makes the behaviour more likely. In everyday language, we might call something a “reward”, but in psychology it is only a reinforcer if behaviour increases.
Positive and negative reinforcement
Before looking at the “properties” of reinforcement, make sure you are clear on two basic forms:
- Positive reinforcement: something pleasant is added after a behaviour, increasing that behaviour.
- Negative reinforcement: something unpleasant is removed after a behaviour, increasing that behaviour.
For example, taking painkillers may be negatively reinforced because the unpleasant pain is removed, making you more likely to take painkillers again in similar circumstances.
Negative reinforcement is not punishment
Negative reinforcement increases behaviour because something unpleasant is removed. Punishment decreases behaviour because it creates an unpleasant consequence or removes something pleasant.
Why reinforcement has “properties”
A reinforcer does not just have one simple effect. Its impact depends on its properties: what kind of reinforcer it is, and how often it is delivered.
The two properties you need for 4.1.4 are:
- Primary versus secondary reinforcement
- Schedules of reinforcement

The big idea
Reinforcement is strongest and most persistent when you understand both what is reinforcing the behaviour and when that reinforcement is delivered.
Property 1: Primary reinforcement
A primary reinforcer is naturally reinforcing because it satisfies a basic biological need or has direct survival value.
Examples include:
- Food
- Water
- Warmth
- Relief from pain
- Sleep or rest
Primary reinforcers do not need to be learned. A hungry animal does not need to be taught that food is reinforcing.
Primary reinforcement
Primary reinforcement occurs when a behaviour is strengthened by a reinforcer that is naturally rewarding, usually because it meets a biological need.
In classic operant conditioning research, B. F. Skinner (1938) used food pellets as a primary reinforcer for rats pressing levers in the “Skinner box”. The food increased lever pressing because it met a biological need.
Property 2: Secondary reinforcement
A secondary reinforcer is not naturally rewarding. Instead, it becomes reinforcing through learning, usually because it has been associated with primary reinforcers or other valued outcomes.
Examples include:
- Money
- Praise
- Grades
- Tokens
- Stickers
- Points in a game
- Social approval
Money is a good example. A £10 note is not biologically useful by itself, but it becomes reinforcing because you have learned that it can be exchanged for primary reinforcers such as food, warmth, and shelter, as well as other desired things.
Secondary reinforcement
Secondary reinforcement occurs when a behaviour is strengthened by a learned reinforcer that has gained value through association with other rewards.
Secondary reinforcement is especially important in human behaviour because many of our reinforcers are social, symbolic, or learned.
Classifying reinforcers in a token economy
A hospital ward uses a token economy. Patients receive tokens for completing daily tasks. Later, they can exchange tokens for snacks, extra leisure time, or privileges.
-
Identify any reinforcers with direct biological value. Snacks may act as primary reinforcers because food is naturally rewarding.
-
Identify any reinforcers that only have value because of learning. Tokens are secondary reinforcers because they are only useful once patients learn they can be exchanged for valued outcomes.
-
Link the reinforcer to behaviour change. If completing tasks is followed by tokens, and the tokens can reliably be exchanged for rewards, patients may be more likely to complete tasks in future.
Quick check
Ask yourself: “Would this be rewarding without learning?” If yes, it is probably primary. If its value depends on association, exchange, status, or approval, it is probably secondary.
Schedules of reinforcement
A schedule of reinforcement is the rule for how often reinforcement is given after a behaviour.
Schedule of reinforcement
A schedule of reinforcement is the pattern or timing by which a behaviour is reinforced.
Schedules matter because they affect:
- How quickly a behaviour is learned
- How often the behaviour is repeated
- How resistant the behaviour is to extinction
Continuous reinforcement
Continuous reinforcement means the behaviour is reinforced every time it occurs.
For example, a rat receives a food pellet every time it presses a lever.
Continuous reinforcement is useful when a behaviour is first being learned because the link between behaviour and consequence is very clear. However, it can also lead to quick extinction if reinforcement stops, because the learner quickly notices the change.
Extinction
Extinction is the gradual weakening and disappearance of a learned behaviour when reinforcement stops.
Partial reinforcement
Partial reinforcement, also called intermittent reinforcement, means the behaviour is reinforced only some of the time.
For example, a person does not win every time they play a fruit machine, but occasional wins may keep them playing.
Partial reinforcement usually makes behaviour more resistant to extinction. This is called the partial reinforcement extinction effect.
Partial reinforcement extinction effect
The partial reinforcement extinction effect is the finding that behaviours reinforced only some of the time are often more resistant to extinction than behaviours reinforced every time.
The logic is simple: if you are used to not being rewarded every time, the absence of reward does not immediately signal that reinforcement has stopped.
The four main partial schedules
Partial schedules are usually described using two distinctions:
- Ratio versus interval
- Fixed versus variable
Ratio schedules
A ratio schedule reinforces behaviour after a number of responses.
- Fixed ratio: reinforcement comes after a set number of responses.
- Variable ratio: reinforcement comes after an unpredictable number of responses.
Interval schedules
An interval schedule reinforces the first response after a period of time has passed.
- Fixed interval: reinforcement becomes available after a set time.
- Variable interval: reinforcement becomes available after an unpredictable time.
Comparing the schedules
| Schedule | How it works | Typical behaviour pattern | Example |
|---|---|---|---|
| Continuous reinforcement | Every correct response is reinforced | Fast learning, but quick extinction | A dog gets a treat every time it sits |
| Fixed ratio | Reinforcement after a set number of responses | High response rate, often with a short pause after reinforcement | A loyalty card gives a free drink after 10 purchases |
| Variable ratio | Reinforcement after an unpredictable number of responses | Very high, steady responding; highly resistant to extinction | Gambling machines or rare item drops in games |
| Fixed interval | Reinforcement for the first response after a set time | Responses increase as the time approaches; “scalloped” pattern | Checking for weekly test results near release time |
| Variable interval | Reinforcement for the first response after unpredictable time periods | Moderate, steady responding | Checking emails or messages when replies arrive unpredictably |
Most persistent schedule
A variable ratio schedule usually produces the strongest resistance to extinction because the learner cannot predict which response will be reinforced.
Identifying a reinforcement schedule
A player defeats enemies in a game. Sometimes a rare item appears after 3 enemies, sometimes after 20, and sometimes after 50. The player keeps playing for long periods.
-
Decide whether reinforcement depends on number of responses or passage of time. The item appears after defeating enemies, so it depends on the number of responses.
-
Decide whether the number is fixed or unpredictable. The number varies each time, so it is variable.
-
Combine the two features. This is a variable ratio schedule.
-
Predict the behavioural effect. The player is likely to show a high response rate and strong resistance to extinction, because each next response might be rewarded.
Why schedules affect extinction
Extinction happens when reinforcement stops. But not all learned behaviours disappear at the same speed.
With continuous reinforcement, the learner expects reinforcement every time. If reinforcement stops, the change is obvious.
With partial reinforcement, non-reward is already normal. The learner has experienced many unrewarded responses before, so they may continue responding for longer.
Predicting resistance to extinction
Two children are rewarded for tidying their rooms. Child A gets praise every single time. Child B gets praise unpredictably, only on some occasions. One month, the praise stops completely.
-
Compare the original schedules. Child A experienced continuous reinforcement, while Child B experienced partial reinforcement.
-
Consider what happens when praise stops. For Child A, the absence of praise is a clear change from the usual pattern. For Child B, no praise is not unusual because it has happened before.
-
Predict persistence. Child B is more likely to continue tidying for longer because partial reinforcement tends to produce greater resistance to extinction.
Evidence and evaluation
AO1: research support
Skinner (1938) provided early experimental evidence for operant conditioning using animals in controlled laboratory settings. Rats or pigeons could be trained to repeat behaviours, such as lever pressing or key pecking, when those behaviours were followed by reinforcement.
Ferster and Skinner (1957) developed detailed work on schedules of reinforcement. Their research showed that different schedules produce different response patterns, such as high steady responding under variable ratio schedules and “scalloped” responding under fixed interval schedules.
AO3: strengths
A major strength is that operant conditioning research is highly scientific. Laboratory studies allow careful control of variables, such as the timing and type of reinforcement. This makes cause and effect easier to establish.
Another strength is real-world application. Reinforcement principles are used in:
- Education, such as praise, points, and reward systems
- Clinical settings, such as token economies
- Animal training
- Behaviour management programmes
- Understanding gambling and game design
AO3: limitations
One limitation is that much early evidence came from animals. Rats and pigeons are useful for controlled research, but human behaviour is more complex. Humans think, plan, form expectations, and respond to social meaning.
Another limitation is that reinforcement can be reductionist. It may explain behaviour mainly in terms of external consequences, while underplaying cognitive factors such as beliefs, motivation, and self-control.
There is also evidence that rewards do not always increase long-term motivation. Deci (1971) found that external rewards can sometimes reduce intrinsic motivation, especially when people originally did an activity because they found it interesting. This is useful AO3 because it shows that reinforcement effects depend on context.
Assuming reinforcement always works
Reinforcement is not magic. Its effect depends on timing, consistency, the learner’s motivation, the value of the reinforcer, and whether the behaviour is already intrinsically rewarding.
Ethics and real-world use
When reinforcement is used with humans, psychologists should consider the BPS Code of Ethics and Conduct (2009). This includes consent, right to withdraw, protection from harm, confidentiality, and debriefing. If deception is used, it must be justified and explained afterwards.
In settings such as schools, prisons, or hospitals, reinforcement programmes can be helpful, but they can also become controlling if people feel manipulated or if access to basic needs is made conditional on behaviour. Token economies, for example, should not remove dignity or deny essential care.
Animal research also raises ethical issues. Skinner’s work involved animals in restricted environments, and modern researchers must consider welfare, deprivation, and harm.
In the exam
-
Define reinforcement precisely: it must increase the likelihood of behaviour being repeated.
-
When applying schedules, first decide whether the rule is based on number of responses or time, then decide whether it is fixed or variable.
-
For AO3, avoid simply saying “it works”. Explain strengths such as control and applications, then balance with limitations such as animal extrapolation, reductionism, ethics, and intrinsic motivation.
Check yourself
- What is the difference between a primary reinforcer and a secondary reinforcer?
- Why does a variable ratio schedule usually produce strong resistance to extinction?
- How could reinforcement be used ethically in a school or clinical setting?
