What you'll learn
- Why psychology needs a scientific community to check new claims.
- How peer review helps validate research before publication.
- The standard format for reporting psychological investigations.
- How to apply this topic in AO1, AO2 and AO3 exam answers.
Why does “validation” matter in psychology?
Psychology aims to produce knowledge about behaviour and mental processes using systematic evidence, not just opinion. A single study can suggest something interesting, but it does not automatically become accepted psychological knowledge.
New knowledge is strengthened when other psychologists can inspect the method, check the analysis, question the interpretation, and attempt to repeat the research.
Scientific community and validation
The scientific community is the network of researchers, journal editors, reviewers, universities, professional bodies and practitioners who produce and evaluate research. Validation means checking whether a new claim is credible, well-supported by evidence, ethically produced, and useful for developing theory or practice.
The big idea
Psychological knowledge becomes more trustworthy when it is publicly reported, critically reviewed, and open to replication by other researchers.
The process below shows how a finding moves from one researcher’s investigation towards wider acceptance.

Peer review
Peer review
Peer review is the process where independent experts in the same area of psychology evaluate a research report before it is published in an academic journal.
A researcher submits a report to a journal. The editor sends it to specialist reviewers, who judge whether the study is good enough to publish. Reviewers usually check:
- whether the research question is clear and worthwhile
- whether the design, sample and procedure are appropriate
- whether ethical guidelines were followed
- whether the data analysis is suitable
- whether the conclusions are justified by the results
- whether the work is original and adds something useful
- whether enough detail is provided for replication
The reviewers then recommend a decision. The paper may be accepted, rejected, or returned for revision and resubmission.
Replication
Replication means repeating a study, often with a new sample or slight changes to the method, to see whether the same pattern of results is found again.
Meta-analysis
A meta-analysis is a statistical review that combines results from several studies on the same topic to estimate the overall strength of an effect.
Why peer review supports valid knowledge
Peer review acts like quality control. It does not prove a study is “true”, but it makes weak claims less likely to enter the published literature unchecked.
For example, if a researcher claims a new therapy reduces anxiety, reviewers may ask whether there was a control group, whether participants were randomly allocated, whether the outcome measure was valid, and whether the correct inferential test was used.
Reviewing a claim about a memory app
A researcher claims that a memory-training app improves recall. They used 12 volunteers, gave them the app for one week, and found a small improvement. They report p=0.08p = 0.08p=0.08 but still claim the app “definitely works”.
-
Evaluate the design. The study lacks a control group, so improvement might be due to practice, motivation, or extra revision rather than the app itself.
-
Check the statistical conclusion. The usual psychology convention is p≤0.05p \leq 0.05p≤0.05 for statistical significance. Since p=0.08p = 0.08p=0.08, the result is not significant at the standard level, so the researcher’s claim is too strong.
-
Assess the sample. A sample of 12 volunteers may be unrepresentative and low in power, meaning a real effect could be missed or an unusual sample could distort the findings.
-
Make a peer-review recommendation. A reviewer might recommend “revise and resubmit”, asking for a control condition, clearer reporting, and a more cautious conclusion.
Evaluating peer review
Strengths
Peer review improves the credibility of psychological knowledge. It can identify confounding variables, inappropriate statistical tests, ethical weaknesses, unclear procedures and over-claimed conclusions.
It also encourages researchers to write in enough detail for others to replicate the work. This links directly to the scientific goals of objectivity, replicability and falsifiability.
Falsifiability
Falsifiability means that a claim is testable in a way that could, in principle, show it to be false.
Limitations
Peer review is not perfect. Reviewers may disagree, miss errors, or be influenced by bias. For example, research from famous universities or well-known researchers may be treated more favourably than equally good work from less prestigious sources.
There is also a risk of publication bias, where journals are more likely to publish significant or exciting results than non-significant findings. This can make an effect look stronger than it really is because unsuccessful replications remain unpublished.
Publication bias
Publication bias occurs when the published literature is unrepresentative because studies with significant, positive or dramatic findings are more likely to be published.
Thinking peer review guarantees truth
Peer review does not prove that a finding is correct. It only means the study has passed expert scrutiny. Strong scientific knowledge usually needs peer review plus replication, transparent reporting and consistency with wider evidence.
Format for reporting psychological investigations
A psychological investigation should be written in a standard format so that other psychologists can understand, evaluate and replicate it.
Research report
A research report is a formal written account of an investigation, usually including the title, abstract, introduction, method, results, discussion, references and appendices.
This standard structure is often called IMRAD: Introduction, Method, Results and Discussion. Psychology reports usually include extra sections before and after this core structure.
The main sections of a psychological report
Title
The title should be concise but informative. It should show the main variables or topic being investigated.
For example, a title such as “The effect of background music on short-term memory recall” is clearer than “Memory experiment”.
Abstract
The abstract is a short summary of the whole investigation. It usually includes the aim, sample, method, key results and conclusion.
You normally write it last, even though it appears near the start.
Introduction
The introduction explains the background research and theory. It should lead logically to the aim and hypotheses.
Hypothesis
A hypothesis is a testable prediction about the relationship between variables. An alternative hypothesis predicts an effect or relationship, while a null hypothesis predicts no effect or relationship.
A report should make clear whether the hypothesis is one-tailed or two-tailed.
- A one-tailed hypothesis predicts the direction of the effect or relationship.
- A two-tailed hypothesis predicts an effect or relationship but not the direction.
One-tailed or two-tailed?
Use a one-tailed hypothesis only when previous theory or evidence gives a clear reason to predict the direction. If the direction is uncertain, use a two-tailed hypothesis.
Method
The method section explains exactly what was done. This is crucial for replication.
It normally includes:
- Design: independent groups, repeated measures, matched pairs, correlation, observation, questionnaire or experiment.
- Participants: sample size, sampling method and relevant characteristics.
- Materials or apparatus: questionnaires, standardised instructions, memory lists, consent forms or equipment.
- Procedure: step-by-step account of what participants did.
- Controls: how extraneous variables were managed.
- Ethics: how the BPS Code of Ethics and Conduct was followed.
Ethical reporting should refer to informed consent, deception, right to withdraw, protection from harm, confidentiality and debriefing. If non-human animals were involved, researchers must also justify animal use and consider welfare, harm reduction and whether alternatives were possible.
Results
The results section presents what was found, without over-interpreting it.
It may include:
- descriptive statistics such as mean, median, mode, range and standard deviation
- tables of raw or summary data
- graphs such as bar charts, histograms, scattergraphs or line graphs
- inferential statistics and the decision about significance
Descriptive and inferential statistics
Descriptive statistics summarise the data you collected. Inferential statistics test whether the results are likely to be due to chance, helping researchers decide whether to reject or retain the null hypothesis.
When reporting inferential statistics, you should state the test used, the observed value, the critical value if using tables, the significance level, whether the test was one- or two-tailed, and the conclusion.
The default significance level in psychology is usually p≤0.05p \leq 0.05p≤0.05.
Choosing and reporting inferential tests
The correct test depends on three main things:
- whether you are testing a difference, relationship, or association
- whether the design uses related or unrelated data
- whether the data are nominal, ordinal or interval
Levels of measurement
Nominal data are categories or frequencies. Ordinal data can be ranked. Interval data use equal numerical units, such as scores on a standardised scale.
Common Eduqas tests include:
- Mann-Whitney U: difference, unrelated design, ordinal data or non-parametric interval data.
- Wilcoxon signed-ranks: difference, related design, ordinal data or non-parametric interval data.
- Spearman’s rho: relationship/correlation, ordinal data or ranked scores.
- Chi-square: association or difference using nominal frequency data.
- Binomial sign test: difference using related nominal data or direction of change.
- Unrelated t-test: difference, unrelated design, interval data, parametric assumptions met.
- Related t-test: difference, related design, interval data, parametric assumptions met.
Reporting a related t-test
A psychologist tests whether the same 20 participants recall more words after using a revision strategy. Recall scores are interval data. The calculated value is t=2.45t = 2.45t=2.45. The critical value for a two-tailed test at p≤0.05p \leq 0.05p≤0.05 with 19 degrees of freedom is 2.093.
-
Identify the design. The same participants are tested before and after the strategy, so the data are related.
-
Identify the level of measurement. Recall is measured as a numerical score with equal units, so the data are treated as interval.
-
Choose the test. A related t-test is appropriate because the study tests a difference using related interval data.
-
Compare observed and critical values. For a t-test, the result is significant when the observed value is equal to or greater than the critical value. Here, 2.45≥2.0932.45 \ge 2.0932.45≥2.093.
-
Write the conclusion. The result is significant at p≤0.05p \leq 0.05p≤0.05, so the null hypothesis is rejected and the strategy appears to affect recall.
Observed versus critical values
For t-tests, chi-square and Spearman’s rho, significance usually requires the observed value to be at least as large as the critical value. For Mann-Whitney U, Wilcoxon signed-ranks and the sign test, significance often requires the observed value to be equal to or smaller than the critical value. Always check the rule for the table you are using.
Discussion
The discussion interprets the findings. It should link back to the hypothesis and previous research, explain what the results suggest, and evaluate the investigation.
A good discussion includes:
- whether the findings support the hypothesis
- possible explanations for the results
- methodological strengths and weaknesses
- issues of reliability and validity
- ethical reflections
- improvements for future research
- real-world applications
This is where AO3 thinking becomes especially important. You might discuss whether the sample was biased, whether demand characteristics affected behaviour, whether the measure was valid, or whether the findings could be useful in education, therapy, health or the workplace.
References and appendices
The references section lists the sources cited in the report. This allows readers to check the background research and prevents plagiarism.
Appendices include supporting materials such as raw data, consent forms, standardised instructions, debrief sheets, questionnaires or calculation details.
Why report format matters
A standard report format makes research transparent. Transparency allows peer reviewers and later researchers to evaluate the study, repeat it, combine it with other evidence, and decide whether it should influence psychological theory or practice.
AO1, AO2 and AO3: how this topic appears in essays
For AO1, describe peer review and the structure of a psychological report accurately.
For AO2, apply these ideas to a scenario. For example, if a student investigation has unclear instructions, no debrief, and no statistical test, you can explain how the report should be improved before it could be taken seriously.
For AO3, evaluate the process. Peer review improves quality and protects science from weak claims, but it may suffer from bias, publication bias, conservatism and human error. Standard reporting improves replicability, but only if researchers are honest and provide enough detail.
In the exam
-
Use the validation chain: report → peer review → publication → replication → theory development.
-
Be specific about report sections: do not just say “write it up”; name sections such as Method, Results and Discussion and explain their purpose.
-
For AO3, balance your answer: peer review is a strength of scientific psychology, but it is not a guarantee of truth because bias, errors and publication bias can still occur.
Check yourself
- What do peer reviewers check before a psychological study is published?
- Why does a standard report format make replication easier?
- How would publication bias affect the conclusions drawn from published research?
