TL;DR
Spaced repetition is one of the best-evidenced study methods in all of health professions education, and manual Anki is its most battle-tested vehicle: a 2026 systematic review in Medical Science Educator found that heavy Anki users outperformed minimal users by 4–13 points on Step 1 across the studies that measured it. So the question worth asking in 2026 is not “does Anki work?” — it does — but “should you still be making the cards by hand?” The best evidence so far, an independent Brown University preprint testing AI-generated Anki cards built from real lecture transcripts, found students performed the same on exams while roughly three-quarters reported meaningful time savings. The catch: card quality depends entirely on how the AI is grounded, and generic chatbots still fabricate references and repeat planted clinical errors at rates that should scare you. The pragmatic answer is a hybrid: keep Anki’s scheduler if you love it, and change how the cards get made — generation grounded in your own lectures, with an export path back into Anki. That is the workflow Neural Consult’s Flashcard Hub was built for.
Why flashcards earned their place — the actual evidence
Flashcards are not a study fad. They sit on top of two of the most replicated findings in cognitive psychology: the spacing effect and retrieval practice. The landmark meta-analysis here is Cepeda and colleagues (2006) in Psychological Bulletin, which synthesized 839 assessments of distributed practice across 317 experiments and found that spacing study out reliably beats cramming — and that the longer you need to remember something, the longer your review intervals should stretch. That last part matters for board exams specifically: you are trying to hold two years of material until a single test date, which is exactly the retention problem spaced algorithms are designed for.
The medical-education data is thinner than the lab data, but it points the same way. The most-cited study is Deng, Gluckstein, and Larsen (2015), which surveyed 72 medical students at Washington University and found that the number of unique Anki cards a student had seen independently predicted their USMLE Step 1 score — roughly one additional point per 1,700 unique cards, after controlling for MCAT score and grades. Notably, board-style practice questions were an even stronger predictor in the same model (about one point per 445 questions), which is worth remembering before you decide flashcards should be your whole plan.
The newer systematic review by Frappa and colleagues (2026) pulled together 11 studies and found a consistent positive association between Anki use and standardized licensing exam performance: high-frequency users outperformed minimal users by 4–13 points, and daily users in one study posted a higher median Step 1 score than non-daily users (238 vs. 233.5, p = 0.039). Course-exam results were more mixed, and the review is careful to note these are observational associations — students who grind thousands of cards a day may simply be more disciplined studiers across the board. Nobody has run the randomized trial that would settle causation. But as study-method evidence goes in this field, spaced retrieval is about as good as it gets, whether you are pointed at USMLE, the NCLEX, the PANCE, or any of the other licensing exams these tools feed.
The case for manual Anki
Let’s give Anki its due, because it has earned it. The desktop and Android apps are free and open source, the scheduler has two decades of refinement behind it, and the ecosystem around it is unmatched. The community-maintained AnKing Step Deck alone contains over 30,000 cards cross-tagged to First Aid, UWorld, Boards & Beyond, Sketchy, and Pathoma, kept current by hundreds of contributors. No AI tool replicates that kind of curated, crowd-tested coverage of a standardized exam, and we are not going to pretend otherwise.
There is also a real argument for making cards yourself. Writing a card forces you to identify what matters in a lecture, compress it, and phrase a retrieval cue — all of which is active processing, not just transcription. Many strong students describe card-making as the moment the material actually gets learned, with the reviews serving as maintenance. If that is you, and you have the hours, manual card-making is a defensible choice on the merits.
And the price is right: free on desktop, with the iOS app running approximately $25 as a one-time purchase (per secondary sources), and the optional AnkiHub sync service around $55 per year.
Where manual decks cost you
The honest ledger has a debit side, and it is mostly measured in hours.
Card-making time is enormous and mostly unexamined. Students routinely spend one to three hours per lecture turning slides into cards — formatting, cloze-deleting, tagging, fixing. Over a preclinical year that is hundreds of hours, and here is the uncomfortable part: the evidence that hand-making cards produces better outcomes than reviewing well-made cards is weak. In the Deng study, the predictor of Step 1 performance was unique cards seen — not cards made. The generation effect from lab psychology is real, but it has never been shown to outweigh the opportunity cost of those hours in a medical curriculum. Most of the community quietly concedes this already: the dominance of premade decks like AnKing is a revealed preference that reviewing beats authoring.
Premade decks solve the time problem but create a relevance problem. AnKing is written to First Aid and the boards — not to your school’s lectures, your professor’s emphasis, or your program’s exams. Nursing, PA, dental, pharmacy, and optometry students feel this hardest, because the mega-decks are overwhelmingly MD/Step-focused. Unsuspending the “right” subset of 30,000 cards to match this week’s lectures is its own skill, and it still leaves your in-house exams only partially covered.
The workload compounds. Anki’s scheduler is unforgiving by design. Miss a week for an away rotation or a family emergency and you return to a four-figure review backlog. None of this means the tool is bad — it means the tool assumes a supply of disciplined hours that many students, especially outside the MD track, do not have.
What the evidence says about AI-generated cards
Until recently, AI flashcards were a promise with no data. That changed with a 2025 study from Brown University’s Warren Alpert Medical School (a preprint, so read with the usual caution — and independent work on a similar tool, not Neural Consult). The researchers built an optimized pipeline that generated Anki flashcards and summaries directly from lecture transcriptions, checked outputs for hallucinations and learning-objective coverage, then deployed the materials across 20 genetics and pharmacology lectures for 143 first-year students. Two findings matter.
First, exam performance showed no significant difference between students who used the AI-generated materials and those who did not. If you were hoping AI cards were a shortcut to higher scores, this is your correction: they are not, at least not yet. Second — and this is the actual value proposition — 74% of students reported time savings from the AI-generated flashcards and 61% from the summaries. Equal learning outcomes for meaningfully less production time is not a small result in a curriculum where time is the binding constraint. The hours you do not spend formatting cloze deletions can go into practice questions — which, recall from Deng, predicted Step 1 performance about four times more efficiently per unit than flashcards did.
Now the caveat that should govern your tool choice: those results came from a carefully engineered, transcript-grounded pipeline with human quality checks — not from pasting slides into a chatbot. The failure modes of ungrounded AI in medicine are well documented. A 2024 study in JMIR found roughly 28.6% of GPT-4-generated medical references were fabricated, and a 2025 study in Communications Medicine found leading models repeated planted clinical errors up to 83% of the time. A flashcard is the worst possible place for an error, because spaced repetition will faithfully burn that error into your long-term memory on an optimized schedule. If you use AI to make cards, the grounding architecture is not a nice-to-have — it is the whole question.
Where Neural Consult fits
Neural Consult’s position here is deliberately not “replace Anki.” It is: keep spaced repetition, change how the cards get made, and keep an exit ramp back to Anki at all times.

The Flashcard Hub generates spaced-repetition decks directly from your own uploaded lectures — the same transcript-grounded approach the Brown study validated, pointed at your school’s actual material rather than a generic syllabus. Because the cards are built from documents you supply, the platform’s grounded-retrieval architecture constrains generation to what your professor actually said, which is the structural answer to the hallucination problem above. And the decks export to Anki, so nothing locks you out of the scheduler and ecosystem you may already trust. If you want the fuller comparison of the two tools as products, we wrote one: Neural Consult vs. Anki.

The cards do not live alone. The same uploaded lectures power the AI Lecture Notebook (chat with your own slides and recordings, the summarization half of what Brown tested), and the material you miss on cards can be drilled as board-style items in the Question Generator — the higher-yield-per-unit practice mode in Deng’s data — or applied in voice-based patient encounters in the Case Simulator.
On the accuracy question, the directly relevant evidence we can point to: in a published test against the official USMLE sample sets, Neural Consult’s engine scored perfectly across Step 1, 2, and 3 (119/119, 120/120, 137/137), and a blinded faculty study rated its AI-generated board questions equal or superior to retired NBME items, with explanations rated 4.4 vs. 3.5 (p < 0.001). Two honest caveats: those benchmarks concern generated questions, not flashcards specifically, and they are medicine benchmarks — validation in other health professions is earlier-stage. More than 75,000 students across MD, DO, PA, nursing, and adjacent programs currently use the platform. For the broader evidence conversation, see Do AI study tools actually work? and, on why a purpose-built tool rather than a chatbot, Neural Consult vs. ChatGPT.
A recommended stack, by studying style
There is no single right answer here — it depends on where your hours currently go.
| Your situation | Spaced repetition | Why |
|---|---|---|
| You already run AnKing daily and it’s working | Keep it. Add AI-generated cards only for school-specific lectures AnKing doesn’t cover, exported into your existing Anki setup. | Don’t fix a working system; the mega-deck’s crowd-tested coverage of the boards is its irreplaceable strength. |
| You make all your cards by hand and you’re drowning | Generate decks from your lectures in Flashcard Hub, review there or export to Anki; spend the reclaimed hours on practice questions. | The Brown data says outcomes hold while production time drops; the Deng data says questions are the higher-yield place for freed-up hours. |
| You’re a nursing / PA / dental / pharmacy / optometry student | Lecture-grounded AI decks as your base, plus any curated deck that exists for your field. | The premade-deck ecosystem is thinnest outside the MD track, so generation from your own curriculum matters most here. |
| You keep abandoning Anki | Auto-generated decks with review built into one platform alongside questions and cases. | The best scheduler is the one you’ll actually open; removing the card-making tax removes the most common quit point. |
The bottom line
The evidence for spaced repetition is real, the evidence for Anki specifically is favorable-but-observational, and the first controlled evidence on AI-generated cards says they trade at par on exam performance while giving you back a large share of your production hours. Nobody’s data — ours included — shows AI cards make you score higher; what they change is what your hours buy. Make the cards by hand if the making is where you learn. But if your card-making has become data entry, the evidence says you can stop: generate from your own lectures with a grounded tool, export to Anki if you want the familiar scheduler, and spend the recovered time on the practice questions that predict scores most efficiently. You can try the full workflow at neuralconsult.com — upload one lecture, generate a deck, and judge the cards against your own by hand.