Anki and medical exam scores: what a systematic review actually shows

Eleven studies suggest Anki often links to better USMLE Step 1 scores when use is frequent — but course exam results are more mixed.

Contents

In one sentence

A systematic review of eleven studies finds that regular Anki use often correlates with higher USMLE Step 1 scores, while evidence for university course exams and Step 2 is more mixed — mostly observational data, not proof for every student.


What the researchers did

Frappa and colleagues searched PubMed, CINAHL, ERIC, PsychINFO, Scopus, Embase, and Web of Science for studies measuring Anki use and exam outcomes in undergraduate medical education. They qualitatively synthesized eleven eligible studies — no single pooled meta-analysis across all effect sizes. Studies differed in how they defined “Anki use,” deck quality, and outcome measures.


What they found

  • Three studies reported a consistent positive link between regular Anki use and USMLE Step 1 performance; high-frequency users scored roughly 4–13 points higher than minimal users in one study, with a suggested dose–response pattern by cards reviewed.
  • University-administered exams: results were mixed — some structured Anki programs showed gains; others found no measurable difference despite positive student perceptions.
  • USMLE Step 2 CK: only one study was available; it found no significant benefit in that sample.
  • Overall evidence is largely observational and heterogeneous — consistent with spaced repetition and retrieval practice theory, but not a guarantee in every context.

What this means for learners and educators

  • Anki is best understood as a delivery system for spacing and retrieval — outcomes depend on deck quality, consistency, and what the exam rewards.
  • For broad foundational exams heavy on factual recall, frequent honest use may pay off; for short in-course tests with narrow timelines, benefits may be harder to see.
  • Faculty should not treat shared student decks as automatically validated — expert review of content still matters, especially for clinical reasoning.

Limitations and what we don't know yet

Eleven studies is a small, varied pool; definitions of “regular use” differ. Most designs are correlational — heavy Anki users may also study more in other ways. Step 2 evidence is almost absent. The review does not replace trials that randomly assign deck-based study or track objective app logs longitudinally.