- What content analysis is and when psychologists use it.
- How to turn written, spoken, or visual material into codes, categories, and themes.
- How content analysis can produce both qualitative and quantitative data.
- How to evaluate content analysis for AO3, including reliability, validity, bias, and ethics.
Psychologists are often interested in what people say, write, post, or show in images and videos. For example, a researcher might want to study how newspapers portray mental illness, how teenagers describe stress in interviews, or how adverts represent gender roles.
This material is called communication data: any recorded human communication that can be analysed, such as interview transcripts, diaries, social media posts, newspaper articles, films, TV programmes, adverts, speeches, or video recordings.
Content analysis
Content analysis is a research technique used to systematically analyse communication data by identifying patterns, meanings, categories, or themes within the material.
Content analysis is especially useful in psychology because it allows researchers to study real-world communication without always needing to create an artificial lab situation.
The big idea
Content analysis is a way of making messy communication data more systematic, so psychologists can draw clearer conclusions from it.
Content analysis often begins with qualitative data.
Qualitative data
Qualitative data is non-numerical data, usually in the form of words, meanings, descriptions, images, or themes.
For example, a participant might say: “I felt watched all the time, so I stopped speaking honestly.” That response gives rich information about feelings and meaning.
However, content analysis can also turn qualitative material into quantitative data.
Quantitative data
Quantitative data is numerical data, such as frequency counts, scores, ratings, or percentages.
For example, a researcher could count how many times a newspaper article uses words linked to danger, such as “threat”, “risk”, or “violent”.
This is why content analysis is sometimes described as a bridge between qualitative and quantitative methods.
A good content analysis follows a clear process: decide what you are studying, collect a sample of communication, create coding categories, code the material, check reliability, then interpret the findings.

The researcher starts with a focused question. For example:
- “How is addiction portrayed in online news articles?”
- “How do students describe exam stress in interviews?”
- “Are male and female athletes described differently in sports commentary?”
A vague question leads to vague coding. A focused question makes the analysis easier to repeat and evaluate.
The sample is the set of communication sources the researcher will analyse. This might be 30 newspaper articles, 20 interviews, 50 adverts, or ten episodes of a TV programme.
The researcher must decide:
- which sources to include
- the time period
- how many items to analyse
- whether the sample is representative
For example, if a researcher only analyses one tabloid newspaper, they should be cautious about generalising to “the media” as a whole.
Coding
Coding means labelling sections of data according to a coding system, so that important features can be counted, compared, or interpreted.
A code is a label attached to a piece of data. For example, in an analysis of interview responses about school stress, codes might include “teacher pressure”, “parent pressure”, “lack of sleep”, and “fear of failure”.
Before coding, the researcher must decide the unit of analysis.
Unit of analysis
The unit of analysis is the size of the communication being coded, such as a word, phrase, sentence, paragraph, image, scene, or whole article.
For example, if studying aggression in films, the unit might be each aggressive act. If studying news articles, the unit might be each headline or each paragraph.
Coding categories
Coding categories are the groups or labels that data are placed into during content analysis.
Good categories should be:
- clear: the coder knows what each category means
- operationalised: defined in observable terms
- mutually exclusive where possible: one item should not fit several categories at once
- relevant to the research question
- complete enough to cover the main material
Operationalisation
Operationalisation means defining a concept in a precise, measurable, or observable way.
For example, “negative language” is too vague. A more operationalised category might be: “words or phrases suggesting danger, blame, disgust, or fear, such as ‘threat’, ‘menace’, ‘contaminated’, or ‘out of control’.”
Using vague categories
Do not write that researchers “look for negative words” without explaining how “negative” is defined. In exams, show that content analysis depends on clear operationalised categories.
Coding interview responses
A researcher wants to analyse interviews with sixth-form students about exam stress.
- The researcher chooses the unit of analysis: each meaningful sentence in the transcript, because one sentence usually contains one main idea.
- They create operationalised categories: “workload pressure” means references to too much work or too little time; “performance fear” means references to failing, grades, or disappointing others; “physical symptoms” means references to sleep, headaches, nausea, or tiredness.
- They apply the categories to a transcript sentence: “I stay up late because I’m scared I’ll let my parents down” is coded as both “physical symptoms” if sleep loss is the focus and “performance fear” because it refers to disappointing parents.
- They tally the codes across all interviews to see which stress themes appear most often.
A theme is a repeated pattern of meaning across the data. Themes are broader than single codes.
Theme
A theme is a recurring idea, meaning, or pattern found across qualitative data.
For example, in interviews about social media use, individual codes might include “checking likes”, “comparing appearance”, and “feeling left out”. A broader theme could be “social media and self-worth”.
Content analysis can be more quantitative, focusing on counting categories, or more qualitative, focusing on interpreting themes and meanings. Many real studies use a mixture of both.
Reliability means consistency. In content analysis, a key issue is whether different researchers would code the same material in the same way.
Inter-rater reliability
Inter-rater reliability is the extent to which two or more independent coders agree when applying the same coding system to the same material.
To improve inter-rater reliability, researchers often use at least two independent coders. They compare their coding and resolve disagreements. If the coders often disagree, the coding categories may be unclear or too subjective.
Reliability shortcut
If an exam asks how content analysis can be made more reliable, mention clear operationalised categories, independent coders, and checking agreement between coders.
Validity means whether the researcher is really measuring or interpreting what they claim to be measuring.
For example, counting the word “crazy” in newspaper articles might seem to measure stigma towards mental illness. But it may not always be valid: the word could be used in unrelated phrases, quotations, jokes, or headlines written for dramatic effect.
So a content analysis can be reliable but not valid. Two coders might consistently count the same words, but those words might not truly represent the psychological concept being studied.
Reliability is not enough
High agreement between coders is useful, but the categories must also capture the real meaning of the communication.
Sometimes content analysis produces nominal data, where data are placed into named categories. For example, articles might be coded as using either “supportive language” or “stigmatising language”.
If the researcher wants to test whether two categories are associated, they may use a Chi-square test.
Nominal data
Nominal data consists of named categories with no numerical order, such as “positive”, “negative”, and “neutral”.
Null hypothesis
A null hypothesis states that there is no difference or association beyond what would be expected by chance.
Choosing a statistical test for coded media data
A researcher compares 40 tabloid articles and 40 broadsheet articles. Each article is coded as either “stigmatising” or “non-stigmatising”.
- The researcher identifies the level of measurement: the data are nominal because each article is placed into a named category.
- The researcher identifies the design: the tabloid and broadsheet articles are independent groups because they are different articles.
- The suitable test is Chi-square because the researcher is testing an association between two nominal variables: type of newspaper and type of language.
- The researcher would look up the critical value for Chi-square using the correct degrees of freedom and the chosen significance level, usually p<0.05p < 0.05p<0.05.
- If the calculated Chi-square value is greater than the critical value, the researcher rejects the null hypothesis; if it is not greater, they fails to reject the null hypothesis.
Suppose a psychologist wants to investigate whether social media posts about body image are more often positive, negative, or neutral.
A strong application would explain that the researcher could:
- collect a sample of posts from a defined platform and time period
- decide whether to code words, captions, comments, images, or whole posts
- create operationalised categories for “positive”, “negative”, and “neutral”
- use two independent coders to apply the coding system
- count the number of posts in each category
- interpret the findings while considering context, sarcasm, and ethical issues
This is better than simply saying “they would analyse the posts”, because it shows you understand the method.
Content analysis often uses naturally occurring communication, such as media articles or online posts. This can increase ecological validity, because the data may reflect real behaviour rather than behaviour produced in a laboratory.
If categories are clearly operationalised, another researcher can repeat the analysis using the same coding system. This improves reliability and makes the method more scientific.
Researchers can analyse many texts, posts, adverts, or videos. They can identify patterns that might not be obvious from reading a few examples casually.
Content analysis can preserve qualitative detail, such as quotes and themes, while also producing quantitative data, such as frequency counts. This can make findings richer and easier to compare.
Even with clear categories, coders may interpret meaning differently. Sarcasm, humour, slang, cultural context, and implied meanings can be hard to code.
For example, the phrase “that exam was sick” could mean very good or very bad depending on context.
Reducing rich communication to counts can lose important detail. Counting how often a word appears does not automatically explain why it was used or how an audience interprets it.
If the sample is narrow or unrepresentative, the results may not generalise. Analysing only one newspaper, one social media platform, or one month of posts may produce a biased picture.
Researchers choose the categories, examples, and interpretations. Their expectations may influence what they notice and what they ignore.
Confusing content analysis with observation
Observation usually involves recording behaviour as it happens. Content analysis involves analysing communication that has been recorded or produced, such as transcripts, articles, posts, or videos.
Content analysis can seem ethically simple because the researcher may use existing material. However, ethical issues still matter.
If the data come from interviews, private messages, diaries, therapy transcripts, or closed online groups, researchers need to consider informed consent, confidentiality, right to withdraw, and protection from harm.
If the data are public, such as newspaper articles or public posts, consent may be less straightforward. Public availability does not always mean people expected their words to be used in research. Researchers should anonymise personal data where possible, avoid exposing vulnerable individuals, and be careful with direct quotations that could be searchable.
Debriefing may be needed if participants directly provided material for the research, especially if deception was involved. Coders may also need support if analysing distressing content, such as trauma accounts or self-harm posts.
Public does not always mean ethically risk-free
Using social media data still raises privacy and confidentiality issues, especially when posts include sensitive information or come from vulnerable groups.
For AO1, define content analysis and outline the process: sample, unit of analysis, coding categories, coding, tallying or theme identification, reliability checks, interpretation.
For AO2, apply those steps to the specific material in the question. If the item is about adverts, mention adverts. If it is about interviews, mention transcripts and quotes.
For AO3, evaluate both the method and how well it was carried out. A content analysis with vague categories and one coder is weaker than one with operationalised categories and high inter-rater reliability.
In the exam
- Start by defining content analysis as a systematic analysis of communication data, not just “looking at texts”.
- When applying it, name the sample, unit of analysis, coding categories, and how reliability would be checked.
- For evaluation, balance strengths such as real-world data and replicability against weaknesses such as subjectivity, loss of meaning, sampling bias, and ethical issues.
Check yourself
- What is the difference between a code, a coding category, and a theme?
- Why might two independent coders improve the reliability of a content analysis?
- How could a content analysis of social media posts raise ethical concerns?