English
September 17, 2026 · 22 min read
The science behind Kastoro's spaced repetition, and how it compares to Anki
Kastoro's flashcards come back shortly before you forget them. This article explains where that sentence comes from: what a century of research supports, what it does not, how our scheduler works under the hood, and where it matches, beats or falls short of Anki's.
1. Forgetting has a shape
In 1885, Hermann Ebbinghaus memorized lists of nonsense syllables and measured how much time he saved when relearning them after 20 minutes, 1 hour, 1 day, up to 31 days [1]. The savings dropped fast at first and slowly afterwards: about 58% after 20 minutes, 34% after a day, 21% after a month. In 2015, Murre and Dros repeated the experiment and obtained a similar curve [2].
Those numbers show up everywhere as “you forget 70% in a day”. That is an overstatement. Each study had only a single participant, the syllables were completely meaningless, and the measure is time saved when relearning, not the percentage remembered. What survives is the shape: the loss is fast at the start and then slows down.
The exact shape matters to a scheduler. Wixted and Ebbesen compared several functions, and the power function described forgetting better than the exponential, across different tasks [3]. Anderson and Tweney objected that averaging several exponential curves can look like a power curve [4]; Wixted and Ebbesen replied by showing the same pattern in individual curves [5]. The power function is the best description available today, not a proven law. It is the shape FSRS uses.
Probability of remembering a card with a stability of 10 days. Both curves pass through 90% on day 10, but the power curve forgets much more slowly afterwards. Computed with the FSRS-6 default parameters (w20 = 0.1542).
2. The two effects everything rests on
Spacing
The meta-analysis by Cepeda and colleagues gathered 839 comparisons from 317 experiments: distributing study over time beats concentrating it, and the ideal gap between sessions grows with how long you want to remember [6]. In a study with more than 1,350 people, the optimal gap was about 20–40% of the delay when the test was a week away, and 5–10% when it was a year away [7]. There is no magic interval; it depends on how long the knowledge has to last. That study used a single review, so it is not a ready-made rule for scheduling dozens of them.
The longest study we know of is the Bahrick family’s: four people, 300 foreign word pairs, tests up to five years later. Thirteen sessions spaced 56 days apart produced retention comparable to 26 sessions spaced 14 days apart [8]. That is only four participants, but the direction matches the rest of the literature: longer gaps produce the same retention with less effort.
Trying to remember
A flashcard is a test, and testing teaches. Roediger and Karpicke showed that rereading wins on a test five minutes later, but trying to recall wins two days and a week later [9]. In the most cited experiment, with 40 Swahili–English pairs, those who kept testing themselves recalled about 80% after a week, compared to 33–36% for those who stopped being tested [10]. That is an extreme laboratory result. Rowland’s meta-analysis, with 159 effect sizes, gives a more realistic figure: g = 0.50, and smaller still in unpublished studies [11]. Moderate, consistent and cheap.
There is an uncomfortable detail: our intuition votes against it. In Kornell’s flashcard study, spacing was better for 90% of participants, and 72% of them believed cramming had worked better [12]. Reviewing what is still fresh feels like mastery. A scheduler exists to make that decision for you.
Difficulty that helps
Bjork and Bjork distinguish between two strengths of a memory: how accessible it is right now and how well it is consolidated. Their thesis is that recalling something that is already hard to reach consolidates it more than recalling something on the tip of your tongue [13]. Pyc and Rawson tested the idea: harder retrievals, as long as they succeeded, gave better final results, with diminishing returns [14]. The authors add a caveat: a difficulty is desirable only when you can overcome it [15]. Reviewing too late is just getting it wrong.
This is where that line comes from. “Shortly before you forget” is the point at which retrieval is hard enough to consolidate and likely enough to succeed.
3. From cardboard boxes to SM-2
In 1967, Pimsleur proposed an illustrative schedule where each review came five times later than the previous one, without fitting it to empirical data [16]. Leitner popularized the card boxes: get it right and the card moves to a box reviewed less often; get it wrong and it goes back to the first one.
In 1987, Piotr Woźniak wrote SM-2 [17]. The rule fits in four lines: the first review comes after 1 day, the second after 6, and from then on each interval is the previous one times an ease factor (EF), which starts at 2.5, goes up or down with your grade and never drops below 1.3. Miss the card and it resets.
Anki adopted SM-2 with documented changes: configurable learning steps instead of 1 and 6 days, four buttons instead of six grades, a bonus for “Easy”, and review lateness factored into the calculation [18]. With that variant, millions of people learned languages and passed medical exams. It works.
SM-2’s limitation is structural. It has no model of memory: it does not estimate your chance of remembering, it only applies a multiplier. Every new card starts with the same factor, and every miss lowers it. Anki’s own manual describes the consequence: repeatedly failing a card can reduce its ease a lot, leading to what some people call “ease hell”, where the card keeps coming back far too often, indefinitely. The manual adds that FSRS does not suffer from this problem [19].
4. FSRS: a model of memory instead of a rule
FSRS (Free Spaced Repetition Scheduler) comes from the work of Jarrett Ye and colleagues, published at KDD in 2022 on 220 million logs from a vocabulary app [20] and extended in 2023 in IEEE TKDE [21]. The conceptual basis is older: the two-component model of long-term memory by Woźniak, Gorzelańczyk and Murakowski [22], the formal cousin of Bjork’s two strengths.
Each card has three variables:
- Retrievability (R): the probability that you remember it now. It falls over time.
- Stability (S): how many days it takes R to fall from 100% to 90%. It is how consolidated the memory is.
- Difficulty (D): from 1 to 10, how strongly this card resists gaining stability.
The FSRS-6 forgetting curve is the power function in the figure above [23]:
R(t, S) = (1 + factor · t / S)−w20, with factor = 0.9−1/w20 − 1
This scaling factor ensures that R is exactly 90% when t = S. When you get a card right, stability grows:
S′ = S · (1 + ew8 · (11 − D) · S−w9 · (ew10 · (1 − R) − 1))
Each term of that formula corresponds to a finding from section 2:
(11 − D): difficult cards consolidate more slowly.S−w9: the more consolidated the memory, the smaller the relative gain. Intervals grow, but less and less quickly.ew10 · (1 − R) − 1: the lower your chance of remembering at review time, the bigger the gain if you succeed. It is desirable difficulty as an equation. With the default weights, a card with stability 20 and difficulty 5 multiplies its stability by about 1.6 if reviewed early (R = 97%), by 3.0 on its due date (R = 90%) and by 7.4 if reviewed late and still remembered (R = 70%).
That correspondence is coherence, not proof. The experiments support the direction of each term; the exact form of the functions was chosen because it fits review data well.
When you miss, stability does not go back to zero. A card with a stability of 100 days and average difficulty falls to about 4 days, and its difficulty goes up. A miss is expensive, but the history is not erased, and difficulty reverts to the mean, so no card stays stuck at the bottom forever.
Finally, the interval is the forgetting curve solved in reverse: the number of days until R falls to the retention you want. With a 90% target, the interval is the stability itself.
5. How Kastoro implements it
Kastoro’s scheduler is FSRS-6 with the 21 published default parameters and a 90% target retention. We checked our implementation against vectors generated by the reference Rust library, the same family of code Anki uses, and against an independent double-precision transcription of the equations.
The four buttons
After seeing the answer you choose between Again, Hard, Good and Easy. The first time you see a card, “Good” schedules it for 2 days, “Hard” for 1 day and “Easy” for 8 days. “Again” brings the card back in 1 minute if it is new, and in 10 minutes if you have seen it before. There is no ladder of learning steps for correct answers: get it right and the card goes straight to the scale of days. Anki’s manual recommends the same approach for FSRS users: short steps that fit in the same day [19].
Answering “Good” on the due date every time, a new card goes through this sequence:
| Review | Kastoro (FSRS-6 default, 90%) | Idealized SM-2 (EF 2.5, first interval of 1 day) |
|---|---|---|
| 1st | 2 days | 1 day |
| 2nd | 11 days | ≈ 3 days |
| 3rd | 46 days | ≈ 6 days |
| 4th | 163 days | ≈ 16 days |
| 5th | 497 days | ≈ 39 days |
In FSRS the multiplier is not constant: 5.5× on the first jump, 3× on the fourth. That does not mean longer intervals are always better. It means that, according to the model, those are the days on which your chance of remembering reaches 90%. If you answer “Hard” on every review, the same card goes 1, 3, 7, 12, 18 days.
What we do differently around the algorithm
- The interval on the button is the interval you actually get. The preview and the write go through the same function, with the same timestamp. If another review arrives from another device between the preview and your tap, the preview is rebuilt.
- The history is the truth; the schedule is derived. What syncs between your devices is the list of answers, which only grows. Stability, difficulty and due date are recomputed from it. Two offline sessions on the same card merge cleanly, no review is discarded, and changing the algorithm in the future requires no migration.
- The same result on iPad and on the Web. The core is the same code, and the scheduler’s math uses its own deterministic arithmetic, verified byte for byte across platforms. Your device and your browser never disagree about when a card is due.
- It works offline. All the computation happens on the device.
- Daily limits that actually count. By default, 20 new cards and 200 reviews per day, counted from the history and respected across sessions. Cards you just missed take priority and are never blocked by the limit.
6. Kastoro and Anki, side by side
| Kastoro | Anki | |
|---|---|---|
| Default algorithm | FSRS-6, always, with nothing to configure. | A variant of SM-2; FSRS is an option in the deck settings. |
| Parameters | The 21 published FSRS-6 default weights. No per-user training. | With FSRS on, an optimizer fits the weights to your history. A real advantage for Anki. |
| Target retention | 90%, fixed today. | Adjustable; the manual recommends 90% as the starting point. |
| Learning steps | Only on a miss: 1 min for a new card, 10 min for a card you have seen. Correct answers go straight to days. | Configurable (1 min and 10 min by default); with FSRS the manual recommends steps shorter than 1 day. |
| Spreading reviews out | No fuzz: cards created together and answered identically come back together. | Automatic fuzz and Easy Days. An advantage for Anki. |
| Source of truth | The history of answers. The schedule is recomputed from it, identically on iPad and on the Web. | Scheduling state stored on the card, plus a review log. |
| What a card is made of | Your notes: text, images, audio, and ink handwritten on the iPad. | HTML note templates, with a huge ecosystem of add-ons and shared decks. |
As far as the memory model goes, Kastoro and Anki with FSRS turned on use the same equations. The honest difference is this: Anki, well configured and with the optimizer run over a few hundred of your reviews, predicts your memory better than Kastoro does today. In exchange, it requires you to know FSRS exists, turn it on, choose a target retention, run the optimizer periodically, and adjust learning steps. Kastoro gives default FSRS-6 to people who have never heard of any of that. Training the weights per person is on our roadmap, and because the schedule is derived from the history, when it arrives it will apply retroactively to all your cards.
7. What the data says, and what it does not
The broadest comparison between schedulers is srs-benchmark, maintained by the group that develops FSRS [24]. The dataset has about 727 million reviews from 10,000 Anki users. For each algorithm it measures how well it predicts whether the person will remember the card. Lower log loss is better.
| Algorithm | Log loss | RMSE (bins) | README version |
|---|---|---|---|
| FSRS-6 optimized per user | 0.346 | 0.065 | August 2026 |
| FSRS-6 with default parameters (what Kastoro uses) | 0.366 | 0.093 | July 2025 |
| Always predicting the person’s average success rate | 0.395 | 0.103 | July 2025 |
| Half-Life Regression (Duolingo, 2016) | 0.469 | 0.128 | August 2026 |
| SM-2 in Anki’s variant | 0.616 | 0.172 | July 2025 |
| Original SM-2 | 0.722 | 0.203 | July 2025 |
Unweighted table, same dataset. The SM-2 and default-parameter rows were removed from the current README; their values come from the July 2025 version.
Read these numbers with four caveats:
- It is prediction, not learning. The benchmark measures calibration: if the algorithm says 90%, do you get it right 90% of the time? Predicting well is a precondition for scheduling well, but nobody has measured in a randomized trial whether students on FSRS learn more than students on SM-2.
- SM-2 was not built for this. It does not produce probabilities; the benchmark added formulas so that it would, and explicitly notes as much. The real gap in prediction is large, but the exact number should be taken with a grain of salt.
- The authors have a vested interest. The code and data are open and anyone can reproduce them, but the benchmark is maintained by the people who created FSRS.
- Default is worse than optimized. 0.366 versus 0.346. And always predicting the person’s average already reaches 0.395. The gain of default FSRS over that floor is real and modest; the gain over SM-2 is large.
You will read that “FSRS reduces reviews by 20 to 30%”. The source is FSRS’s own documentation, which notes that the number comes from simulation [23]. The 12.6% in the 2022 paper and the 17% in the 2023 paper are also simulations on the authors’ memory model. They are reasonable indications, not measurements on students, and that is why we do not promise a percentage.
The best classroom evidence that personalized scheduling helps comes from Lindsey and colleagues: in one school, model-personalized review improved retention by 16.5% over massed study and by 10.0% over one-size-fits-all spacing [25]. It is a single school, with a model other than FSRS. It shows that the general idea works outside the lab.
8. What we do not claim
On the iPad, Kastoro cards can be written by hand. It would be tempting to say that handwriting makes you remember more. The literature does not readily support that. The most recent meta-analysis of lecture notes found a small advantage for handwriting on paper (g = 0.25) [26]. The most cited study favoring longhand [27] was not confirmed in a direct replication [28]. None of these studies tested a stylus on a tablet.
What stands on firmer ground is the generation effect: information you produce yourself is remembered better than information you simply read, with a mean effect of 0.40 across 86 studies [29]. It is an argument for making your own cards from your notes, with a keyboard or a pen, instead of only downloading a ready-made deck. Ready-made decks still have their place for covering content. They just do not replace the work of formulating the question.
Nor do we claim that flashcards raise your grades. The studies of medical students that associate Anki use with higher scores are correlational and small. The randomized trial that exists shows what the theory predicts: self-testing beat rereading after one week, and the difference disappeared at six months without further reviews [30]. The effect depends on continuing to review, and keeping that on track is a scheduler’s job.
9. How to use it well, in any app
- Answer honestly. If you did not remember, it is “Again”, not “Hard”. Anki’s manual warns that using “Hard” for misses inflates every interval [19], and the same holds here.
- Show up every day, even briefly. The model assumes you review close to the due date. Delays will not break the schedule, but facing a backlog is discouraging.
- Distrust the feeling of mastery. The review that seems unnecessary and the one that seems too hard tend to be the ones that work.
- Make small cards, from your notes. One question, one answer.
Cards that come back shortly before you forget.
Make yours from your own notes, on the iPad or in the browser.
References
- Ebbinghaus, H. (1885). Über das Gedächtnis. Duncker & Humblot. Translated by Ruger and Bussenius (1913): psychclassics.yorku.ca.
- Murre, J. M. J., & Dros, J. (2015). Replication and analysis of Ebbinghaus’ forgetting curve. PLOS ONE, 10(7), e0120644. doi:10.1371/journal.pone.0120644
- Wixted, J. T., & Ebbesen, E. B. (1991). On the form of forgetting. Psychological Science, 2(6), 409–415. doi:10.1111/j.1467-9280.1991.tb00175.x
- Anderson, R. B., & Tweney, R. D. (1997). Artifactual power curves in forgetting. Memory & Cognition, 25(5), 724–730. doi:10.3758/BF03211315
- Wixted, J. T., & Ebbesen, E. B. (1997). Genuine power curves in forgetting. Memory & Cognition, 25(5), 731–739. doi:10.3758/BF03211316
- Cepeda, N. J., Pashler, H., Vul, E., Wixted, J. T., & Rohrer, D. (2006). Distributed practice in verbal recall tasks. Psychological Bulletin, 132(3), 354–380. doi:10.1037/0033-2909.132.3.354
- Cepeda, N. J., Vul, E., Rohrer, D., Wixted, J. T., & Pashler, H. (2008). Spacing effects in learning: A temporal ridgeline of optimal retention. Psychological Science, 19(11), 1095–1102. doi:10.1111/j.1467-9280.2008.02209.x
- Bahrick, H. P., Bahrick, L. E., Bahrick, A. S., & Bahrick, P. E. (1993). Maintenance of foreign language vocabulary and the spacing effect. Psychological Science, 4(5), 316–321. doi:10.1111/j.1467-9280.1993.tb00571.x
- Roediger, H. L., & Karpicke, J. D. (2006). Test-enhanced learning. Psychological Science, 17(3), 249–255. doi:10.1111/j.1467-9280.2006.01693.x
- Karpicke, J. D., & Roediger, H. L. (2008). The critical importance of retrieval for learning. Science, 319(5865), 966–968. doi:10.1126/science.1152408
- Rowland, C. A. (2014). The effect of testing versus restudy on retention: A meta-analytic review of the testing effect. Psychological Bulletin, 140(6), 1432–1463. doi:10.1037/a0037559
- Kornell, N. (2009). Optimising learning using flashcards: Spacing is more effective than cramming. Applied Cognitive Psychology, 23(9), 1297–1317. doi:10.1002/acp.1537
- Bjork, R. A., & Bjork, E. L. (1992). A new theory of disuse and an old theory of stimulus fluctuation. In Healy, Kosslyn & Shiffrin (Eds.), From Learning Processes to Cognitive Processes, vol. 2, 35–67. Erlbaum.
- Pyc, M. A., & Rawson, K. A. (2009). Testing the retrieval effort hypothesis. Journal of Memory and Language, 60(4), 437–447.
- Bjork, E. L., & Bjork, R. A. (2011). Making things hard on yourself, but in a good way. In Gernsbacher et al. (Eds.), Psychology and the Real World, 56–64. Worth.
- Pimsleur, P. (1967). A memory schedule. Modern Language Journal, 51(2), 73–75.
- Woźniak, P. A. Algorithm SM-2, excerpted from Optimization of learning (1990). super-memory.com
- Anki FAQ. What spaced repetition algorithm does Anki use? faqs.ankiweb.net. Accessed 17 September 2026.
- Anki Manual. Deck Options. docs.ankiweb.net. Accessed 17 September 2026.
- Ye, J., Su, J., & Cao, Y. (2022). A stochastic shortest path algorithm for optimizing spaced repetition scheduling. KDD ’22, 4381–4390. doi:10.1145/3534678.3539081
- Su, J., Ye, J., Nie, L., Cao, Y., & Chen, Y. (2023). Optimizing spaced repetition schedule by capturing the dynamics of memory. IEEE Transactions on Knowledge and Data Engineering, 35(10), 10085–10097. doi:10.1109/TKDE.2023.3251721
- Woźniak, P. A., Gorzelańczyk, E. J., & Murakowski, J. A. (1995). Two components of long-term memory. Acta Neurobiologiae Experimentalis, 55(4), 301–305. doi:10.55782/ane-1995-1090
- Open Spaced Repetition. The Algorithm; ABC of FSRS. github.com/open-spaced-repetition/awesome-fsrs. Accessed 17 September 2026.
- Open Spaced Repetition. srs-benchmark. github.com/open-spaced-repetition/srs-benchmark. README of 29 August 2026 and of 24 July 2025 (commit 45f61b2).
- Lindsey, R. V., Shroyer, J. D., Pashler, H., & Mozer, M. C. (2014). Improving students’ long-term knowledge retention through personalized review. Psychological Science, 25(3), 639–647. doi:10.1177/0956797613504302
- Flanigan, A. E., Wheeler, J., Colliot, T., Lu, J., & Kiewra, K. A. (2024). Typed versus handwritten lecture notes and college student achievement: A meta-analysis. Educational Psychology Review, 36, 78. doi:10.1007/s10648-024-09914-w
- Mueller, P. A., & Oppenheimer, D. M. (2014). The pen is mightier than the keyboard. Psychological Science, 25(6), 1159–1168. doi:10.1177/0956797614524581
- Morehead, K., Dunlosky, J., & Rawson, K. A. (2019). How much mightier is the pen than the keyboard for note-taking? Educational Psychology Review, 31, 753–780. doi:10.1007/s10648-019-09468-2
- Bertsch, S., Pesta, B. J., Wiscott, R., & McDaniel, M. A. (2007). The generation effect: A meta-analytic review. Memory & Cognition, 35(2), 201–210. doi:10.3758/BF03193441
- Schmidmaier, R., Ebersbach, R., Schiller, M., Hege, I., Holzer, M., & Fischer, M. R. (2011). Using electronic flashcards to promote learning in medical students: Retesting versus restudying. Medical Education, 45(11), 1101–1110. doi:10.1111/j.1365-2923.2011.04043.x