Memory & Learning: How to Actually Make Things Stick
The three-stage memory model, why working memory is the bottleneck, the forgetting curve, and the two techniques with the strongest evidence in all of psychology — spacing and retrieval practice.
Almost everyone studies wrong. Re-reading and highlighting — the two most popular techniques — are among the least effective known, while the two most effective ones feel bad while you do them. This doc covers how memory actually works, why forgetting is a feature, and the handful of techniques with mountain-sized evidence behind them.
The standard model: three stores
The classic Atkinson–Shiffrin architecture, still the right first approximation:
- Sensory memory buffers raw perception for a fraction of a second. Whatever you don’t attend to never existed, as far as memory is concerned.
- Working memory is where thinking happens — and it holds only about 4 chunks (the folklore “7±2” was optimistic; modern estimates are 3–5). This is the hard constraint behind interface design, teaching, and why you can’t follow a 9-step verbal instruction.
- Long-term memory is effectively unlimited, but the problem was never storage — it’s retrieval. “Forgotten” usually means inaccessible, not erased, which is why a cue (a smell, a song) can resurrect a memory you’d have sworn was gone.
Chunking is the exploit for the tiny middle box: CBINASAPIN is ten items,
but CBI · NASA · PIN is three — same information, restructured around
knowledge you already have. Expertise largely is chunking: a chess master sees
“a Sicilian defence,” not 32 piece positions, freeing working memory for actual
thinking.
Forgetting: the curve and why it’s a feature
Ebbinghaus (1885, testing himself on nonsense syllables — psychology’s first great self-experiment) found memory decays fast then slow: huge losses in the first day, a long stable tail after. Roughly half of new, meaningless material is gone within hours if untouched.
But forgetting isn’t a bug. A memory that never faded would drown you in noise — you need last week’s parking spot to fade so today’s can win. Forgetting is the system betting that unused information is unneeded. Which points directly at the fix: prove the information is needed, repeatedly, at the moments it’s about to fade.
The two techniques that actually work
Decades of research, hundreds of studies, and the ranking is stable. The winners:
1. Retrieval practice (the testing effect)
Pulling a memory out strengthens it far more than putting it in again. Closing the book and forcing recall — flashcards, blank-page brain dumps, practice problems — beats re-reading by large margins in study after study. The brutal part: retrieval feels worse. Re-reading feels smooth and familiar (“I know this!”), but that fluency is System 1 mistaking recognition for recall. Struggling to retrieve feels like failure and is precisely the event that builds the memory.
2. Spacing (distributed practice)
Five hours across two weeks beats five hours in one night — same total time. Cramming works for tomorrow’s exam and evaporates by next month. Each spaced review catches the memory mid-fade, and the effort of that harder retrieval is what multiplies retention. The optimal schedule expands: review at ~1 day, ~3 days, ~1 week, ~1 month. This is exactly what Anki and every spaced-repetition system automates — each card gets its own expanding schedule based on your recall history.
The compounding combo (for anything you must retain):
1. Learn it → immediately self-test, book closed
2. Next day → retrieve again (expect struggle; that's the mechanism)
3. Then at ~3d, 1w, 1m — or just let Anki schedule it
4. Got it wrong? The interval resets short. Right? It expands.
Also genuinely useful
- Elaboration — asking why and how does this connect to what I know. Memory is a web; unconnected facts have no retrieval paths.
- Interleaving — mixing problem types (ABCABC not AAABBB). Feels harder and messier, tests better, because you also learn to recognize which approach fits — which is what exams and life actually demand.
- Concrete examples & dual coding — pair every abstraction with an example and, where possible, a picture (two retrieval paths beat one — it’s why these docs have diagrams).
Overrated, to be clear
Re-reading and highlighting produce fluency, not learning. And “learning styles” (visual/auditory/kinesthetic learners) is one of psychology’s most persistent myths — matching teaching to a person’s declared style shows no benefit in controlled tests. Everyone learns better from good multimodal material; nobody has a magic modality.
Memory is reconstructive (the unsettling part)
Memory is not a video recorder — every recall is a reconstruction, rebuilt from fragments plus your current knowledge, expectations, and mood. Elizabeth Loftus’s experiments showed that wording alone rewrites memories: witnesses who heard “how fast were the cars going when they smashed into each other?” later remembered broken glass that never existed; those who heard “hit” didn’t. Entire false childhood memories can be implanted in a substantial minority of people through repeated suggestive interviews.
Practical consequences: eyewitness confidence is a poor guide to accuracy (a leading cause of wrongful convictions), your vivid “flashbulb” memories of big events drift dramatically over years while feeling rock-solid, and — see hindsight bias — your memory of what you predicted quietly edits itself to match what happened.
Takeaways
- Working memory is the bottleneck: ~4 chunks. Chunk, or overload.
- Forgetting is fast-then-slow and adaptive — fight it at the fade points.
- Test yourself and space it out. The techniques that feel worst work best; fluency is a lie System 1 tells you.
- Skip highlighting marathons and learning-styles quizzes.
- Trust memories less than they ask you to — especially vivid, confident ones.
The same machinery drives behaviour change — which is where motivation and habits come in.