What you'll learn
- What coding means in content analysis.
- How psychologists create units of analysis, coding categories and a coding frame.
- How coding can turn qualitative material into data that can be counted or compared.
- How to evaluate coding using reliability, validity, sampling and ethical issues.
The big picture: what is content analysis?
Psychologists often study communication: interviews, therapy transcripts, diaries, adverts, newspapers, films, school policies, online posts or TV programmes. These data are usually rich and detailed, but they can be messy. Content analysis is a way to make this material more systematic.
Content analysis
Content analysis is a research technique used to systematically study communication material by identifying, coding and analysing patterns in its content.
For example, a researcher might analyse newspaper articles to see how often mental illness is linked with danger, or analyse children’s TV shows to see how often male and female characters are shown taking leadership roles.
Coding
Coding means labelling parts of the communication material according to categories, so that the content can be counted, compared or interpreted consistently.
The key idea is that the researcher turns raw material — words, images, actions, phrases — into organised data.

Why coding matters
Coding is the bridge between rich qualitative material and systematic analysis. It helps the researcher move from “I noticed a pattern” to “I applied clear rules and found this pattern.”
Before coding: decide what you are analysing
Communication material
Communication material means the source being analysed. This could be written, spoken, visual or digital.
Examples include:
- interview transcripts
- newspaper headlines
- social media comments
- adverts
- song lyrics
- films or TV programmes
- clinical case notes
- political speeches
A good content analysis starts with a clear research question. For example: “How are teenagers portrayed in newspaper headlines?” is clearer than “What are newspapers like?”
Sampling the material
Researchers rarely analyse every possible item. Instead, they select a sample, meaning a smaller set of material chosen to represent the wider group of interest.
If the sample is biased, the coding may be neat but the conclusion may still be weak. For instance, analysing only crime stories about teenagers would probably exaggerate negative portrayals.
Forgetting the sample
A coding system can be reliable but still produce a misleading conclusion if the material selected is unrepresentative.
The unit of analysis
Before coding, the researcher must decide exactly what counts as one “thing” to code.
Unit of analysis
The unit of analysis is the specific part of the material that will be coded, such as a word, phrase, sentence, image, speaker turn, article, scene or behaviour.
The unit depends on the research question. If you are analysing gender stereotypes in adverts, the unit might be each advert, each character, or each action performed by a character. These choices can lead to different results.
Examples of units of analysis
If the material is a newspaper article, the unit could be:
- each word, such as “dangerous”, “violent” or “vulnerable”
- each sentence
- each whole article
- each reference to a social group
- each image or caption
If the material is a TV programme, the unit could be:
- each character
- each scene
- each speaking turn
- each example of aggressive behaviour
Choosing the unit
A useful unit is small enough to be coded precisely, but large enough to keep the meaning. A single word is easy to count, but a sentence or scene may capture context better.
Coding categories
Once the unit is chosen, the researcher creates coding categories.
Coding category
A coding category is a label used to classify a unit of analysis, such as “positive portrayal”, “negative portrayal”, “physical aggression” or “supportive comment”.
Good coding categories should be operationalised. This means the researcher clearly defines exactly what counts as an example.
Operationalisation
Operationalisation means defining a concept in a precise, measurable way so that different researchers know how to identify it.
For example, “aggression” is too vague by itself. A more operationalised category would be: “physical aggression: any deliberate hitting, pushing, kicking or throwing an object at another person.”
Mutually exclusive and exhaustive categories
Categories are usually stronger if they are mutually exclusive, meaning one unit cannot fit into more than one category at the same time. They should also be as exhaustive as possible, meaning they cover all important possibilities.
For example, if coding the tone of headlines, the categories might be:
- positive
- negative
- neutral or mixed
Without “neutral or mixed”, coders may be forced to make unreliable guesses.
Creating a coding frame for headlines
A researcher wants to analyse how teenagers are portrayed in newspaper headlines.
-
The researcher chooses the headline as the unit of analysis because the research question focuses on first impressions created by news reporting.
-
The researcher creates three operationalised categories: positive means the headline presents teenagers as helpful, successful or admirable; negative means it presents them as criminal, lazy, threatening or irresponsible; neutral or mixed means the headline is factual or contains both positive and negative elements.
-
The researcher checks whether the categories overlap. A headline such as “Teen volunteers praised after flood rescue” fits positive, while “Teen gang blamed for station attack” fits negative. A headline such as “Teenagers respond to new school rules” fits neutral.
-
The researcher writes these rules into a coding frame so another coder can apply the same categories without needing to guess the researcher’s intentions.
The coding frame
Coding frame
A coding frame is a document that lists the coding categories, their definitions, decision rules and examples, so coders apply the system consistently.
A coding frame might include:
- the research question
- the unit of analysis
- each category name
- a definition of each category
- examples and non-examples
- rules for unclear cases
This is important because content analysis should be systematic. Another researcher should be able to understand how the coding was done and, ideally, repeat the procedure.
Coding frame
The coding frame is like the instruction manual for the analysis. It reduces guesswork and helps make the findings more replicable.
Manifest and latent content
There are two useful levels of content that researchers may code.
Manifest content
Manifest content is the obvious, surface-level content that is directly present, such as specific words, phrases, images or actions.
For example, counting how many times a newspaper article uses the word “violent” is coding manifest content.
Latent content
Latent content is the underlying meaning, implication or theme in the material, which requires interpretation.
For example, a TV advert might never say “women belong in the home”, but it may repeatedly show women doing housework while men make financial decisions. Coding that gender stereotype involves latent content.
Manifest coding is usually easier to replicate because it is more objective. Latent coding may be richer and more meaningful, but it can be more subjective.
Only counting words
Content analysis is not always just word-counting. Sometimes the most important pattern is in the meaning, tone, role or implication of the material.
Quantitative and qualitative uses of coding
Coding can produce quantitative data, meaning numerical data such as frequencies or percentages. For example, a researcher might count how many articles are positive, negative or neutral.
Coding can also support qualitative analysis, meaning analysis focused on meanings, themes and detailed examples. For instance, the researcher might describe recurring themes in how anxiety is discussed in student blogs, using quotations to support each theme.
In A-Level Psychology, content analysis is often useful because it can combine both:
- numerical summaries, such as how often a category appears
- rich examples, such as quotations or descriptions of scenes
Link to statistical testing
Coded data are often nominal data, meaning data placed into named categories with no natural order, such as “positive”, “negative” and “neutral”.
If a researcher wants to test whether category frequencies differ between independent groups, a Chi-square test may be appropriate. For example, they could compare whether tabloid and broadsheet newspapers differ in how often they portray teenagers negatively. The researcher would compare the calculated Chi-square value with the critical value from a table, usually using p<0.05p < 0.05p<0.05 as the default significance level. If the calculated value is equal to or larger than the critical value, the null hypothesis is rejected.
Reliability in coding
Reliability
Reliability means consistency. In content analysis, reliable coding means the same material would be coded in the same way by the same coder over time or by different coders.
A key form is inter-rater reliability.
Inter-rater reliability
Inter-rater reliability is the extent to which two or more independent coders agree when applying the same coding frame to the same material.
Researchers can improve inter-rater reliability by:
- training coders
- using clear operational definitions
- piloting the coding frame on a small sample
- revising vague categories
- using examples and non-examples
- discussing ambiguous cases before final coding
Improving inter-rater reliability
Two coders analyse social media posts about exam stress using the categories “practical support”, “emotional support” and “negative pressure”.
-
They independently code the same small sample and find that they often disagree on posts like “You’ll be fine if you actually revise”, because one coder sees it as practical support and the other sees it as negative pressure.
-
They identify the source of disagreement: the category definitions do not explain how to code advice that includes criticism.
-
They revise the coding frame so “practical support” includes specific advice without blame, while “negative pressure” includes criticism, threat, shame or blame.
-
They code a new pilot sample using the revised rules. If agreement improves, the coding frame is more reliable.
Reliability is not the same as validity
Two coders can agree perfectly and still be measuring the wrong thing. A coding frame may be consistent but not valid if it fails to capture the psychological concept properly.
Validity in coding
Validity
Validity means whether a method measures what it claims to measure.
In content analysis, validity depends on whether the coding categories genuinely capture the concept being studied.
For example, if a researcher claims to measure “stigma towards mental illness” but only counts the word “mad”, they may miss more subtle forms of stigma, such as linking mental illness with danger, unpredictability or helplessness.
Ways to improve validity include:
- basing categories on previous research or theory
- using both manifest and latent categories where appropriate
- including examples from the material to support interpretations
- checking whether the categories fit the research question
- avoiding categories that are too broad or too narrow
AO3 evaluation: strengths of coding in content analysis
It is systematic and replicable
A clear coding frame makes content analysis more objective than simply reading material and giving an impression. Other researchers can inspect the categories and repeat the coding.
It can analyse real-world material
Content analysis often uses naturally occurring communication, such as newspapers, adverts or online posts. This can improve ecological validity, because the material was not created artificially for a laboratory study.
It can handle large amounts of data
Coding allows researchers to summarise large volumes of material. For example, hundreds of articles can be coded into categories and compared.
It can combine depth and numbers
Content analysis can provide frequency counts while still using quotations or examples to illustrate meaning. This is useful in psychology because behaviour and communication often need both measurement and interpretation.
AO3 evaluation: limitations of coding in content analysis
Researcher bias can affect categories
The researcher chooses the categories, definitions and examples. If these are shaped by expectations, the findings may reflect the researcher’s assumptions rather than the material itself.
Meaning can be lost
When rich communication is reduced to categories, context may disappear. A sarcastic comment might be coded as “positive” if the coder focuses only on the words and misses the tone.
Categories may be too rigid
If categories are decided before looking at the material, unexpected themes may be missed. However, if categories are developed during analysis, the process may become more subjective unless carefully documented.
Sampling problems can weaken conclusions
A content analysis is only as good as the material selected. A study of “media portrayals” based on one newspaper, one week or one platform may not generalise well.
AO3 balance
For evaluation, pair each strength with a condition. For example: content analysis is systematic if the coding frame is clear, and valid if the categories genuinely capture the concept.
Ethical considerations
Content analysis can seem ethically simple because researchers often use existing material. But ethical issues still matter.
If the material is public, such as newspaper articles, consent is usually less of a concern. However, if the material involves private diaries, therapy transcripts, private messages or identifiable social media posts, researchers must consider informed consent, confidentiality and the right to withdraw.
Researchers should anonymise names, usernames, locations and other identifying details where needed. They should also protect people from harm, especially when analysing sensitive topics such as mental illness, trauma, offending or self-harm.
If participants have provided material directly, they should usually be told how it will be used and debriefed afterwards if the full purpose could not be explained beforehand.
Assuming online means ethical
Just because something is online does not automatically mean it is ethically safe to quote or analyse without care. Think about privacy, identifiability and possible harm.
How to apply this in a scenario
If you are given a research scenario, look for these decisions:
- What is the communication material?
- What is the unit of analysis?
- What are the coding categories?
- Are the categories operationalised?
- Are they mutually exclusive and sufficiently exhaustive?
- How could the researcher check inter-rater reliability?
- What ethical issues apply?
In the exam
-
For AO1, define coding clearly: it is the process of applying labels or categories to communication material in content analysis.
-
For AO2, apply your answer to the exact material in the scenario, such as articles, posts, transcripts or adverts. Name a sensible unit of analysis and give examples of categories.
-
For AO3, evaluate both reliability and validity. Clear categories improve reliability, but subjective interpretation, poor sampling and loss of context can reduce validity.
-
If ethics are relevant, mention consent, confidentiality, right to withdraw, protection from harm and debrief — especially for private or identifiable material.
Check yourself
- What is the difference between a unit of analysis and a coding category?
- Why might latent content be harder to code reliably than manifest content?
- How could a researcher improve inter-rater reliability in a content analysis?
