Skip to main content

Literature review

How to Conduct a Systematic Literature Review: A Step by Step Guide

A systematic literature review is not a longer narrative review. It is a method: a pre-specified, transparent procedure for finding, filtering, appraising, and synthesizing every study relevant to one focused question, documented well enough that another team could rerun it and land in the same place. This guide walks the full pipeline, from framing the question to reporting the PRISMA flow, with a worked micro-example and the numbers that make each step auditable.

11 min read · Updated August 21, 2026

The defining property of a systematic review is reproducibility. Where a narrative review reflects an author's reading and judgement, a systematic review commits to a protocol before the first search runs, then executes it so faithfully that the search string, the databases, the dates, the screening decisions, and the exclusion reasons are all recoverable. That discipline is what lets a systematic review sit at the top of the evidence hierarchy, and it is the same provenance thinking that a good research operations platform applies at every stage. If you have not yet mapped where a review fits in your process, start with the modern research workflow and return here for the deep pass.

Systematic vs narrative vs scoping: pick the right instrument

Before you commit months to a review, be sure the systematic form is the one your question needs. The three common review types answer different questions and demand different rigor.

  • Narrative review: an expert synthesis of a topic, selective by design, useful for context and theory-building. It does not claim exhaustive coverage and rarely reports a search strategy, so a reader cannot check what was left out.
  • Scoping review: maps the breadth of a field, what has been studied, with what methods, and where the gaps are. It follows a structured search and screening process (often reported with the PRISMA-ScR extension) but usually stops short of formal quality appraisal or pooled estimates.
  • Systematic review: answers one focused question by locating all eligible studies, appraising their risk of bias, and synthesizing them. When the included studies are similar enough, it may add a meta-analysis that statistically pools their results.

A quick test: if you cannot state your question in a single sentence with a defined population and outcome, you probably want a scoping review first. If you can, and you intend to weigh the evidence rather than merely catalogue it, you want a systematic review.

Frame a focused, answerable question

Everything downstream inherits the precision of your question: a vague question produces a vague search, an unmanageable screening pile, and a synthesis that says little. Two frameworks structure the question so its components map directly onto search terms and eligibility rules.

PICO for quantitative and clinical questions

PICO decomposes a question into Population, Intervention, Comparison, and Outcome (sometimes a fifth term, Study design). Our worked example throughout this guide: does spaced retrieval practice improve long-term retention in undergraduate STEM students compared with massed restudy? The Population is undergraduate STEM students, the Intervention is spaced retrieval practice, the Comparison is massed restudy, and the Outcome is retention on a delayed test at least one week later. Each slot becomes a block in the search string and a line in the eligibility criteria.

SPIDER for qualitative and mixed-methods questions

PICO assumes an intervention and a measurable outcome, which fits qualitative work poorly. SPIDER (Sample, Phenomenon of Interest, Design, Evaluation, Research type) is built for questions about experience and meaning, for example how first-generation graduate students experience imposter feelings in their first year. Match the framework to your evidence base rather than forcing an interview study into an intervention shape.

Register the protocol first

Write and register your protocol before you search. PROSPERO accepts systematic review protocols in health and social care; for other fields, deposit a timestamped protocol in the Open Science Framework. Registration separates a planned analysis from a post-hoc story, and it is the strongest defence against the charge that you cherry-picked.

Write a reproducible search strategy

The search is the part reviewers scrutinize hardest, because a review can only be as complete as the search that fed it. Build it deliberately, then record it verbatim.

Controlled vocabulary and free-text terms

Databases index articles with controlled vocabularies: MeSH (Medical Subject Headings) in PubMed, Emtree in Embase, thesaurus descriptors in PsycINFO. Controlled terms catch articles regardless of the authors' wording, so "spaced practice" is found even when the paper says "distributed practice". Combine those subject headings with free-text keywords to catch recent papers not yet indexed. Truncation (retriev* matches retrieval, retrieves, retrieving) and proximity operators widen the net without listing every variant.

Boolean logic

Assemble the query as blocks joined by AND, with synonyms inside each block joined by OR. For the worked example one block might read: ("spaced practice" OR "distributed practice" OR "spaced retrieval" OR "retrieval practice") AND ("long-term retention" OR "delayed test" OR "durable learning") AND (undergraduate* OR "college student*"). AND narrows, OR broadens, and NOT is dangerous because it silently discards borderline records, so avoid it.

Choosing databases

  • Scopus and Web of Science: broad, multidisciplinary, with cited-reference data that powers citation chasing.
  • PubMed / MEDLINE: essential for biomedical and health questions, with full MeSH indexing.
  • PsycINFO, ERIC, CINAHL: discipline-specific depth for psychology, education, and nursing respectively.
  • Google Scholar: enormous recall but no controlled vocabulary and limited export, so use it to supplement (screening perhaps the first 100 to 200 results), not as a primary database.

Search at least two or three databases; a single source misses roughly a third of eligible studies in most fields. Then run supplementary methods: backward citation chasing (scanning the reference lists of included papers), forward chasing (finding papers that cite them), and a check of grey literature such as theses, preprints, and conference proceedings to counter publication bias. Record the exact string, interface, and date for every database, because databases update daily and an undated search cannot be reproduced.

Define inclusion and exclusion criteria before screening

Eligibility criteria are the rules that decide, without further judgement calls, whether a record belongs. Fix them before you see the results so that a large or inconvenient study cannot tempt you into rewriting the rules around it. Derive each criterion straight from the question components.

  1. 1Population: undergraduate students in STEM fields. Exclude K-12 samples, graduate-only samples, and clinical populations.
  2. 2Intervention and comparison: a spaced or distributed practice condition against a massed or restudy control. Exclude single-condition studies with no comparison.
  3. 3Outcome: retention assessed on a delayed test at least seven days after practice. Exclude immediate-only post-tests.
  4. 4Design: randomized or quasi-experimental. Exclude purely correlational and case reports.
  5. 5Report characteristics: peer-reviewed or preprinted, in English, published from 2000 onward, with extractable quantitative outcomes.

Every full-text exclusion must cite one of these rules, because those reasons become the annotations on the PRISMA diagram. If you find yourself inventing a new reason mid-screen, stop and amend the protocol explicitly rather than quietly.

Run the PRISMA flow with real counts

PRISMA (Preferred Reporting Items for Systematic Reviews and Meta-Analyses) is the reporting standard, and its flow diagram is the spine of a review. It moves records through four stages, and at each transition you report how many records entered, how many left, and why. Here is the worked example populated with counts.

Step 1Identification

1,284 records from four databases plus 37 from citation chasing and grey literature, for 1,321 total. Remove 429 duplicates, leaving 892 unique records.

Step 2Screening

Screen all 892 titles and abstracts against the eligibility criteria. Exclude 764 as clearly irrelevant, carrying 128 forward.

Step 3Eligibility

Retrieve and read all 128 full texts. Exclude 96 with recorded reasons: 41 wrong population, 29 no delayed test, 18 no comparison, 8 not retrievable in full.

Step 4Inclusion

32 studies meet every criterion and enter the synthesis. Of those, 24 report enough statistics (means, standard deviations, sample sizes) to enter a meta-analysis.

These numbers are not decoration. A reader who sees 1,321 identified narrowing to 32 included can judge whether your search was too tight or your criteria too loose, and reconstruct where the evidence base came from. PRISMA 2020, the current update, refined the diagram and added a 27-item reporting checklist worth completing as you write.

Screen in duplicate and measure agreement

A single screener is a single point of failure: fatigue and unconscious preference leak in. The standard is dual independent screening, where two reviewers assess every record blind to each other, then reconcile disagreements by discussion or a third adjudicator.

Cohen's kappa and inter-rater reliability

Raw percentage agreement flatters you, because two screeners who both reject 90 percent of records agree most of the time by chance alone. Cohen's kappa corrects for that. It runs from 0 (no better than chance) to 1 (perfect), and the common reading is that 0.61 to 0.80 is substantial and above 0.80 almost perfect. Suppose our two reviewers screened the 892 abstracts and agreed on 831; kappa against the expected chance agreement is about 0.79. That is high enough to proceed, and low enough to justify a reconciliation meeting on the 61 records where they diverged. Report the value in your methods; it is direct evidence that screening was not arbitrary. For the reading discipline that makes full-text screening faster, see reading and annotating research papers efficiently.

Pilot the criteria on a sample

Before full screening, both reviewers should independently screen the same 50 records and compare. A low pilot kappa usually means the criteria are ambiguous, not that a reviewer is careless. Refine the eligibility rules until they are sharp, then screen the full set. Fixing the instrument early is far cheaper than reconciling hundreds of disagreements later.

Extract data and appraise each study

Build a structured extraction table

Once studies are included, pull the same fields from each into a single extraction table, one row per study. Designing that table well is what turns 32 papers into analyzable data. Typical columns:

  • Citation key and DOI, so every row is traceable to its source.
  • Sample: size, field, institution type, mean age.
  • Intervention detail: spacing gap, number of sessions, materials.
  • Comparison condition and how the control spent equivalent time.
  • Outcome: retention interval, test format, and the effect statistic (mean difference, Cohen's d, or odds ratio).
  • Risk-of-bias judgements for each domain.

Extract in duplicate for at least a subset and check for transcription errors, because a single mis-keyed standard deviation can distort a pooled estimate. Keep the table linked to the underlying PDFs rather than retyped in isolation, so that when a reviewer questions a number you can jump from the cell straight to the highlighted passage. Holding that link is what it means to keep evidence connected to its sources, the difference between a defensible extraction and a spreadsheet nobody trusts.

Appraise quality and risk of bias

Not all included studies deserve equal weight. Quality appraisal, or risk-of-bias assessment, evaluates how far each study's design protects its findings from systematic error. Use a validated tool matched to the design: Cochrane RoB 2 for randomized trials, ROBINS-I for non-randomized studies, the Newcastle-Ottawa Scale for observational work, and the JBI checklists for qualitative and prevalence studies.

Each tool walks specific domains, for randomized trials the randomization process, deviations from intended interventions, missing outcome data, measurement of the outcome, and selective reporting. Two reviewers rate each domain as low, some concerns, or high risk, then justify the call. Crucially, you do not simply discard high-risk studies; you carry the judgement into the synthesis and test whether excluding them changes the conclusion, a sensitivity analysis.

A systematic review does not become trustworthy because it is long. It becomes trustworthy because every decision it made is written down and could be checked.

Synthesize: narrative or meta-analytic

Synthesis is where the included studies become an answer. There are two broad routes, and the choice is dictated by how similar the studies are, not by preference.

Narrative synthesis

When studies differ too much in population, intervention, or outcome to combine numerically, you synthesize narratively: group findings by theme or moderator, describe the direction and consistency of effects, and explain the heterogeneity rather than papering over it. Structured methods such as SWiM (Synthesis Without Meta-analysis) keep this transparent by requiring an explicit rule for how studies were grouped and compared.

Meta-analysis

When outcomes are comparable, meta-analysis pools them into a single weighted estimate. Each study contributes a standardized effect size (Cohen's d for continuous outcomes, log odds ratio for binary ones), weighted by its precision, usually under a random-effects model that lets the true effect vary across studies. You report the pooled effect with a confidence interval, quantify inconsistency with the I-squared statistic, and inspect a funnel plot for the asymmetry that signals publication bias. In the worked example, pooling the 24 quantitative studies might yield a moderate benefit of spacing, d of about 0.45, with I-squared of 60 percent indicating substantial but explainable heterogeneity, which you probe by subgroup and meta-regression. Turning those numbers into a defensible position is the same craft covered in building an evidence base you can defend.

Manage hundreds of references without losing provenance

A real review handles a thousand records and a hundred included studies, each with a PDF, a screening decision, extraction fields, and a bias rating. The failure mode is drift: a reference manager that disagrees with the screening spreadsheet, a citation whose DOI no longer resolves, an included study nobody can map back to the search that found it. Discipline here is not tidiness; it is the integrity of the review.

  • Give every record a stable identifier (DOI where available) and a single canonical BibTeX entry, so citations resolve consistently from screening through to the manuscript.
  • Deduplicate on identifiers before screening, not by eyeballing titles, which misses formatting variants.
  • Keep the audit chain intact: each included study should trace from its row in the extraction table back to the exact database search, date, and query that surfaced it.
  • Version the protocol and record any amendment with a date and reason, so the final method matches what you actually did.

For keys, styles, and BibTeX at scale, see citation styles explained, and for the archiving that lets others rerun your search, the guide to research reproducibility.

One connected chain, from search to submission

A systematic review is the purest test of connected provenance in research: it fails the moment any link, from a database query to a screening decision to a pooled effect to a cited claim, cannot be traced back to where it came from. That is what Research Woven is built to hold. Every source keeps its highlights, every extraction stays anchored to the passage it came from, every claim in the write-up points back through its evidence to the study and page that support it, and the counts on your PRISMA diagram stay in sync with the records they describe. The human still frames the question, judges the risk of bias, and writes the argument; the platform operates everything around it so nothing drifts. Run your next review on a connected research operations platform and the audit trail writes itself as you work.

Frequently asked questions

How long does a systematic literature review take?
A rigorous systematic review typically takes six months to over a year with a small team. The search and screening of hundreds to thousands of records is the slowest phase, especially with dual independent screening. Scoping reviews are usually faster because they skip formal quality appraisal and meta-analysis.
What is the difference between a systematic review and a meta-analysis?
A systematic review is the whole method: framing a question, searching, screening, appraising, and synthesizing. A meta-analysis is one optional synthesis technique within it that statistically pools comparable results into a single weighted estimate. Every meta-analysis rests on a systematic review, but many systematic reviews synthesize narratively instead because their studies are too heterogeneous to pool.
How many databases should I search for a systematic review?
Search at least two to three databases, chosen to cover your field, for example Scopus or Web of Science for breadth plus a discipline-specific source such as PubMed, PsycINFO, or ERIC. A single database misses roughly a third of eligible studies. Supplement with citation chasing and a grey-literature check to counter publication bias.
What is a good Cohen's kappa for screening agreement?
Cohen's kappa corrects agreement between two screeners for chance. Values of 0.61 to 0.80 are read as substantial agreement and above 0.80 as almost perfect. A low kappa usually signals ambiguous eligibility criteria rather than a careless reviewer, so refine the criteria on a pilot sample before full screening.
Do I have to exclude studies with a high risk of bias?
Not automatically. You record the risk-of-bias judgement for each study and carry it into the synthesis, then run a sensitivity analysis that recomputes the result with high-risk studies removed. If the conclusion holds either way, it is robust; if it flips, the evidence base depends on weak studies and you report that honestly.
What does PRISMA actually require?
PRISMA is a reporting standard, not a search method. It asks you to report your review transparently: the flow diagram with record counts at identification, screening, eligibility, and inclusion, plus a 27-item checklist covering the question, search strategy, eligibility criteria, and synthesis. Following it makes your review reproducible and reviewable.

Bring this into your own research

Research Woven connects your sources, highlights, notes, evidence, and manuscript in one maintained chain, so provenance and citations are computed for you rather than pieced together by hand.

Open Research Woven

Keep reading