

0 / 2 embers
0 / 3000 xp
click for more info
Complete a lesson to start your streak
click for more info
Still calibrating
click for more info
Not enough gems
Cost: 6 gems
1: Manual Evaluation
incomplete
2: Golden Dataset
incomplete
3: Precision Metrics
incomplete
4: Recall Metrics
incomplete
5: F1 Score
incomplete
6: Error Analysis
incomplete
7: LLM Evaluation
incomplete
Back
ctrl+,
Next
ctrl+.
This lesson's interactive features are locked, please to keep using them
Precision measures the quality of search results. Now let's look at recall – a metric that answers a different question.
Precision asks: "How much of what you found is relevant?"
Recall asks: "How much of what's relevant did you find?"
Recall measures completeness. It tells you what percentage of all relevant documents you actually retrieved.
recall = relevant_retrieved / total_relevant
Imagine searching for "family friendly bear movies":
Recall = 4 relevant retrieved / 8 total relevant = 50%
You found half of the relevant movies, but missed the other half!
Higher recall often means lower precision. Retrieving more documents increases the chance of including irrelevant ones.
The key is finding the right balance for your use case. This is as much of a product question as it is a technical one.
Update the evaluation script to also print recall@k scores for each test case.
k=10
- Query: cute british bear marmalade
- Precision@10: 0.1000
- Recall@10: 1.0000
- Retrieved: Paddington, The Duchess, The Bear, The Great Bear, Care Bears Movie II: A New Generation, Care Bears Nutcracker Suite, Bugs Bunny and the Three Bears, The Edge, An Unfinished Life, The Care Bears Adventure in Wonderland
- Relevant: Paddington
- Query: talking teddy bear comedy
- Precision@10: 0.2000
- Recall@10: 1.0000
- Retrieved: Ted, Ted 2, The Bear, Littleman, Arsenic and Old Lace, Testament, Aqua Teen Hunger Force Colon Movie Film for Theaters, Stand by Me, Desierto, Memento
- Relevant: Ted 2, Ted
Run and submit the CLI tests.