
Evaluating retrieval quality
Why measure retrieval separately from the answer?
Because an answer that was never retrieved cannot be generated. Recall@k, the share of the answering documents that a retriever puts in its top k results, is the ceiling on everything downstream, and no prompt, model or reranker lifts it. So when a RAG answer comes out wrong, the first thing to check is whether the right document reached the context at all. That is a retrieval measurement, and it is separate from anything the generation model did with what it was given.
The course has made two notable claims so far:
Lesson 05 said hybrid search beats either half of it, and
Lesson 09 said a smaller chunk retrieves better than a large one.
However, neither claim has been tested, which is why this lesson builds the harness that tests them, and reports what it finds even where that disagrees with the lesson making the claim.
Step 1 - start a server
The corpus, the analyzer and the two indexes are lesson 05's, unchanged: ten help-centre articles for a fictional online store, a BM25 full-text index and an HNSW vector index. This lesson adds no retrieval feature, and measures the ones you already have.
An in-memory instance in its own terminal is authenticated and throwaway:
surreal start --user root --pass secret
surreal import --endpoint http://localhost:8000 \
--user root --pass secret --ns ai --db eval schema.surql
surreal import --endpoint http://localhost:8000 \
--user root --pass secret --ns ai --db eval seed.surqlThree files go in: schema.surql for the corpus, the golden set, three retrievers and the metrics, seed.surql for ten articles and eight labelled questions, and queries.surql for the evaluation.
Step 2 - the golden set is a table
A golden set is a list of questions with the documents that answer them, labelled by a human. Most teams keep it in a spreadsheet or a JSON file. Keep it in the database instead and the labels can be typed:
DEFINE TABLE OVERWRITE question SCHEMAFULL;
DEFINE FIELD OVERWRITE text ON question TYPE string;
DEFINE FIELD OVERWRITE embedding ON question TYPE array<float>
ASSERT $value.len() = 4;
DEFINE FIELD OVERWRITE relevant ON question TYPE array<record<article>>
ASSERT $value.len() > 0;array<record<article>> means every label has to point at an article rather than at some other table, which an evaluation set kept in a spreadsheet cannot check at all. The type does not check that the article still exists, though: delete one and the label sits there pointing at nothing. Adding REFERENCE ON DELETE REJECT to the field closes that gap, refusing to delete an article any question still depends on, and ON DELETE UNSET is the other option, which drops the deleted article from the label instead.
The type still allows an empty array, which is what the ASSERT is for. A question nobody labelled scores zero for every retriever and quietly drags every average down.
The eight questions are phrased the way a customer would phrase them, not the way the articles are written:
CREATE question:signin SET
text = "I can't get into my account",
embedding = [0.95, 0.0, 0.0, 0.1],
relevant = [article:password_reset, article:account_locked, article:enable_2fa];
CREATE question:rate_limit SET
text = "429 too many requests",
embedding = [0.25, 0.25, 0.25, 0.25],
relevant = [article:api_rate_limits];Two of those numbers are the whole experiment. The sign-in question shares no vocabulary with any of its three answers, so it can only be found semantically. The 429 question is a bare error code, whose embedding is a vague direction rather than a topic, so its vector is near-uniform on purpose. Lesson 05 makes the same modelling choice for the same reason.
Reading the set back gives you the review artefact, which is the other reason to keep it here (three of the eight records):
SELECT id, text, relevant.title AS answers FROM question ORDER BY id;[
{ id: question:money_back, text: 'how do I get my money back',
answers: ['Request a refund', 'Cancel or pause a subscription'] },
{ id: question:rate_limit, text: '429 too many requests',
answers: ['API rate limits and 429 errors'] },
{ id: question:signin, text: "I can't get into my account",
answers: ['Reset a forgotten password', 'Why is my account locked?', 'Set up two-factor authentication'] }
]Step 3 - one shape for every retriever
A metric needs nothing from a retriever except record IDs, best first. Reduce every candidate to that shape and they become comparable:
-- Semantic only, over the HNSW index.
DEFINE FUNCTION OVERWRITE fn::retrieve_vec($qvec: array<float>, $k: int) -> array {
RETURN (SELECT id, vector::distance::knn() AS d FROM article
WHERE embedding <|10, 40|> $qvec ORDER BY d ASC LIMIT $k).id;
};
-- Lexical only. '@1,OR@' matches any term and exposes BM25 as search::score(1).
DEFINE FUNCTION OVERWRITE fn::retrieve_bm25($q: string, $k: int) -> array {
RETURN (SELECT id, search::score(1) AS s FROM article
WHERE content @1,OR@ $q ORDER BY s DESC LIMIT $k).id;
};
-- Both, fused on rank with the Reciprocal Rank Fusion from lesson 05.
DEFINE FUNCTION OVERWRITE fn::retrieve_hybrid(
$q: string, $qvec: array<float>, $k: int
) -> array {
LET $bm25 = SELECT id, search::score(1) AS s FROM article
WHERE content @1,OR@ $q ORDER BY s DESC LIMIT 10;
LET $vec = SELECT id, vector::distance::knn() AS d FROM article
WHERE embedding <|10, 40|> $qvec ORDER BY d ASC;
RETURN search::rrf([$bm25, $vec], $k, 60).id;
};The KNN operator's K has to be a literal, so fn::retrieve_vec asks for the whole corpus and lets the caller's $k do the cutting. On a real corpus you would set that literal to your largest candidate pool.
Here are the three on the 429 question:
LET $qn = question:rate_limit;
RETURN fn::retrieve_bm25($qn.text, 3);
RETURN fn::retrieve_vec($qn.embedding, 3);
RETURN fn::retrieve_hybrid($qn.text, $qn.embedding, 3);[article:api_rate_limits, article:account_locked, article:refund_request]
[article:return_item, article:cancel_subscription, article:enable_2fa]
[article:refund_request, article:account_locked, article:api_rate_limits]The labelled answer is article:api_rate_limits. BM25 puts it first, the vector retriever misses it entirely, and hybrid recovers it in third place.
Step 4 - the metrics are functions
There are three, and they fit in a screen. recall@k asks how much of the answer you found:
-- recall@k: of the articles that answer this question, how many are in the top k?
DEFINE FUNCTION OVERWRITE fn::recall_at($retrieved: array, $relevant: array, $k: int) -> float {
LET $top = $retrieved[0..math::min([$k, $retrieved.len()])];
RETURN <float> $top.intersect($relevant).len() / $relevant.len();
};precision@k asks what fraction of the context you spent was worth spending:
-- precision@k: of the articles sent to the model, how many belonged there?
-- Every irrelevant one is context budget spent on nothing.
DEFINE FUNCTION OVERWRITE fn::precision_at($retrieved: array, $relevant: array, $k: int) -> float {
LET $top = $retrieved[0..math::min([$k, $retrieved.len()])];
RETURN IF $top.len() = 0 { 0f } ELSE { <float> $top.intersect($relevant).len() / $top.len() };
};Reciprocal rank asks where the first correct hit landed, and averaged over a set it is the mean reciprocal rank, or MRR:
-- Reciprocal rank: 1 / the position of the first relevant hit, or 0 if there is none.
DEFINE FUNCTION OVERWRITE fn::rr($retrieved: array, $relevant: array) -> float {
LET $first = $retrieved.map(|$id| $relevant CONTAINS $id).find_index(true);
RETURN IF $first = NONE { 0f } ELSE { 1f / ($first + 1) };
};The math::min([$k, $retrieved.len()]) is there because a range index past the end of an array evaluates to NONE rather than to the short array, so a retriever that returned fewer than k results would silently score zero without that clamp. This is the kind of bug an eval harness must avoid.
This is what the semantic retriever returns for the 429 question:
LET $got = fn::retrieve_vec($qn.embedding, 3);
RETURN { recall_at_3: fn::recall_at($got, $qn.relevant, 3),
precision_at_3: fn::precision_at($got, $qn.relevant, 3),
rr: fn::rr($got, $qn.relevant) };{ recall_at_3: 0f, precision_at_3: 0f, rr: 0f }Step 5 - run the set
fn::evaluate($k) runs all eight questions against all three retrievers and returns the means:
RETURN [1, 3, 5].map(|$k| fn::evaluate($k));[
{ k: 1, bm25: 0.438f, vector: 0.667f, hybrid: 0.479f },
{ k: 3, bm25: 0.771f, vector: 0.875f, hybrid: 0.833f },
{ k: 5, bm25: 0.833f, vector: 0.875f, hybrid: 0.958f }
]That table does not say what lesson 05 led you to expect, so let's take a look at why that is the case.
At k = 3, the hybrid retriever loses to vector search alone, 0.833 against 0.875. At k = 5 it wins, 0.958 against 0.875. Fusion needs room: with only three slots, the lexical half spends one of them on a candidate the semantic half would not have chosen, and on this corpus that trade is a loss. Raise k and the same trade becomes a gain.
MRR tells a consistent story at k = 5:
[{ mrr_bm25: 0.646f, mrr_vec: 0.875f, mrr_hybrid: 0.76f }]Hybrid finds more of the answers than vector search and puts the first one lower down. If your prompt takes five documents, that is fine. If only the top one is ever used, it is not. The metric you choose has to match the shape of what you build.
Step 6 - the mean hides what you need to see
The same run gives one record per question:
SELECT text, relevant.len() AS labels,
math::round(fn::recall_at(fn::retrieve_bm25(text, 3), relevant, 3) * 1000) / 1000 AS bm25,
math::round(fn::recall_at(fn::retrieve_vec(embedding, 3), relevant, 3) * 1000) / 1000 AS vec,
math::round(fn::recall_at(fn::retrieve_hybrid(text, embedding, 3), relevant, 3) * 1000) / 1000 AS hybrid
FROM question;| question | labels | bm25 | vector | hybrid |
|---|---|---|---|---|
| my credit card was declined | 1 | 1.0 | 1.0 | 1.0 |
| how do I get my money back | 2 | 0.0 | 1.0 | 0.0 |
| where is my parcel | 1 | 1.0 | 1.0 | 1.0 |
| 429 too many requests | 1 | 1.0 | 0.0 | 1.0 |
| I want to send this item back | 2 | 0.5 | 1.0 | 1.0 |
| I can't get into my account | 3 | 0.667 | 1.0 | 0.667 |
| stop charging me every month | 1 | 1.0 | 1.0 | 1.0 |
| how do I verify a webhook payload | 1 | 1.0 | 1.0 | 1.0 |
The 0.042 the mean moved between vector and hybrid is two questions cancelling out. Hybrid fixed the 429 question completely, from 0 to 1.0, which is the case lesson 05 was built for. It also broke "how do I get my money back" completely, from 1.0 to 0.
A mean would have hidden both. Always keep the per-question records: the aggregate tells you whether the change is an improvement, and the individual records tell you what to fix.
Step 7 - diagnose the one that broke
LET $bad = question:money_back;
RETURN { question: $bad.text, labelled: $bad.relevant.title,
vector: fn::retrieve_vec($bad.embedding, 3).title,
hybrid: fn::retrieve_hybrid($bad.text, $bad.embedding, 3).title };{
question: 'how do I get my money back',
labelled: ['Request a refund', 'Cancel or pause a subscription'],
vector: ['Request a refund', 'Cancel or pause a subscription', 'Update your payment method'],
hybrid: ['Return an item', 'Track your order', 'API rate limits and 429 errors']
}Vector search got it exactly right. Hybrid returned three articles, none of them labelled. The lexical list is where to look:
SELECT title, math::round(search::score(1) * 100) / 100 AS bm25
FROM article WHERE content @1,OR@ "how do I get my money back" ORDER BY bm25 DESC;[
{ title: 'Return an item', bm25: 2.46f },
{ title: 'Track your order', bm25: 1.27f },
{ title: 'API rate limits and 429 errors', bm25: 1.15f }
]There it is. "Return an item" contains "send something back", and @1,OR@ matches any term, so a question about money matches an article about parcels on one word. Three weak matches come back, all of them wrong and all of them ranked.
RRF fuses on rank and throws the scores away, which is what makes it robust when both lists are decent. The cost is that it cannot tell corroboration from coincidence: a rank-1 candidate in a list of pure noise carries the same weight as a rank-1 candidate in a good list. Two of these noise articles also appear low in the vector list, so they collect points from both sides and outvote the right answers.
Step 8 - measure the fix
The hypothesis: BM25 candidates that score 2.46 on a six-word question are noise, and fusing them pushes better candidates out of the top k. So require a lexical candidate to clear a floor before it is fused at all:
-- Both, fused on rank, with a floor on the lexical list: a BM25 candidate has to
-- score at least $floor to be fused at all.
DEFINE FUNCTION OVERWRITE fn::retrieve_hybrid_floor(
$q: string, $qvec: array<float>, $k: int, $floor: float
) -> array {
LET $bm25 = SELECT id, search::score(1) AS s FROM article
WHERE content @1,OR@ $q ORDER BY s DESC LIMIT 10;
LET $vec = SELECT id, vector::distance::knn() AS d FROM article
WHERE embedding <|10, 40|> $qvec ORDER BY d ASC;
RETURN search::rrf([$bm25.filter(|$r| $r.s >= $floor), $vec], $k, 60).id;
};Now it is a measurement rather than an opinion:
RETURN [0f, 1.5f, 2.5f, 4f].map(|$floor| fn::evaluate_floor(3, $floor));[
{ k: 3, floor: 0f, recall: 0.833f },
{ k: 3, floor: 1.5f, recall: 1f },
{ k: 3, floor: 2.5f, recall: 1f },
{ k: 3, floor: 4f, recall: 1f }
]Recall@3 goes from 0.833 to 1.0, beating both halves and the unfloored fusion. At k = 5 it goes from 0.958 to 1.0.
The uncomfortable part is also the most useful thing in this lesson. Every floor from 1.5 to 4.0 scores identically. Eight questions cannot tell them apart, so picking 2.5 because it is in the middle is a coin flip with a decimal point on it. And the floor was chosen by looking at the failure it fixes, on the same set used to score it, which is how you overfit an evaluation set. The number to report is from questions the change has not seen.
Step 9 - keep the run
An evaluation you run once tells you where you stand today. An evaluation you keep tells you when something has got worse, because each run can be compared against the ones before it:
DEFINE TABLE OVERWRITE eval_run SCHEMAFULL;
DEFINE FIELD OVERWRITE at ON eval_run TYPE datetime DEFAULT time::now();
DEFINE FIELD OVERWRITE retriever ON eval_run TYPE string;
DEFINE FIELD OVERWRITE k ON eval_run TYPE int;
DEFINE FIELD OVERWRITE recall ON eval_run TYPE float;
DEFINE FIELD OVERWRITE questions ON eval_run TYPE int;LET $n = (SELECT count() FROM question GROUP ALL)[0].count;
CREATE eval_run SET retriever = 'hybrid+floor', k = 3,
recall = fn::evaluate_floor(3, 1.5f).recall, questions = $n;
SELECT retriever, k, recall, questions FROM eval_run ORDER BY recall DESC;[
{ retriever: 'hybrid+floor', k: 3, recall: 1f, questions: 8 },
{ retriever: 'vector', k: 3, recall: 0.875f, questions: 8 },
{ retriever: 'hybrid', k: 3, recall: 0.833f, questions: 8 },
{ retriever: 'bm25', k: 3, recall: 0.771f, questions: 8 }
]The history lives in the same database as the corpus it measures, so a run is one CREATE and a trend is one SELECT. If you run queries.surql twice you get two sets of records. The questions field is there so a number from a 40-question set is never compared with a number from an 8-question one.
From here it can run on its own. Wrap the run in a function and call it whenever documents are ingested, so a drop in the numbers appears next to the data that caused it. Nothing has to leave the database, and what is measured cannot drift away from what does the measuring.
How big does a golden set need to be?
It needs to be bigger than eight, and the reason is resolution rather than any particular threshold. Recall@k is not binary once a question has more than one relevant document - the run in step 6 scored 0.5 and 0.667 on questions labelled with two and three documents - so the step a single question contributes to the mean is 1 / (labels × question_count), not a flat 1 / question_count. A retrieval change on "I can't get into my account" (3 labels) moves the eight-question mean by 1 / (3 × 8), about 4.17 percentage points; the same kind of change on a single-label question like "429 too many requests" moves it the full 1/8, or 12.5 points. Step 8 is what the coarse end of that range looks like: four different floors scored identically because the gaps between them were smaller than even the single-label step. Adding questions shrinks that coarse step - still roughly 1 / question_count - and questions with more labels sharpen the mean's resolution further on top of that, so the set does not cross a line at some number, it just gets finer.
The general advice in statistics that a sample should be at least 30 comes from a different question, which is roughly how many samples make the distribution of a mean close enough to normal to put a confidence interval around it. It is a convention rather than a boundary, and it is answering "how certain am I of this number" where the paragraph above answers "how small a difference can I see at all". Both happen to point at a few dozen questions.
Some working numbers:
30 to 50 questions is the smallest set where a mean starts to mean anything, and one question still moves it 2 to 3 points.
100 to 200 if you want to slice by category or question type and read the slices.
Every question that ever failed in production belongs in the set permanently. That is the part that stops the same regression shipping twice.
Where questions come from matters more than how many there are. Take them from support tickets, search logs and the questions users actually asked. Questions written by the person who wrote the corpus test the corpus against itself, and they always score better than reality.
Label what actually answers the question, not everything on the topic. Every extra label lowers recall for a retriever that was right, and the metric then rewards a retriever that returns broad, vague context.
What this evaluation cannot tell you
Four things sit outside what it measures:
Whether these numbers transfer. The corpus is ten documents with hand-written 4-dimensional vectors. That makes semantic search unrealistically strong: the question vectors were written to point at the topics the articles cover, with none of the noise a real model produces. Vector search scoring 0.875 here says nothing about your corpus. Run this harness on your data, with your model, before you rely on any of these figures.
Whether the answer was right. Recall@k is a ceiling, not a guarantee. A model handed the correct document can still contradict it. Faithfulness and answer quality need their own evaluation, with the generated answer in the loop.
Whether it is fast enough, or cheap enough. Neither appears in any of these metrics. The floored hybrid retriever runs two index scans and a fusion for every query.
Whether the labels are right. The golden set is a human artefact and it goes stale as the corpus changes. While array<record<article>> stops a label pointing at a non-article, nothing stops a label pointing at an article that has since been rewritten to say something else.
From this set to yours
- 1.
Replace the corpus with your documents, and
DIMENSION 4with your model's dimensionality. - 2.
Write 30 questions from your support log, and label them by hand.
- 3.
Keep
fn::retrieve_*, the three metrics andfn::evaluateas they are. They only assume a ranked list of record IDs. - 4.
Put the run in CI, on a fixed
kthat matches how many documents your prompt actually takes.
Then use it on the decision from lesson 09. Chunk your corpus two ways into two tables, point the retriever at each, and let recall@k choose. That gives you a number for the choice, which is more than any chunking guide can offer.
THE PLATFORM
Everything an application and its agents know. Five surfaces, one engine.
Database
Document, graph, vector, time-series and relational in one engine.

Agent Memory
What an agent learns, with its source and its time, in the same engine.

Cloud
Managed clusters in the regions you choose, scaled on demand.

Studio
Query, explore and design the schema from the browser.

MCP
Every model that speaks MCP reaches the database and the memory directly.

IN PRODUCTION
Trusted at scale. Samsung, Nvidia, Verizon, Tencent, and Walmart run on SurrealDB.
14,000+
Developers building on SurrealDB Cloud
4M+
Developers building on SurrealDB worldwide
FROM THE TEAMS
SurrealDB gives us a foundation where we can unify semantic search, knowledge graphs, and AI-driven decision making without stitching together multiple systems. Collapsing responsibility into SurrealDB has become our default engineering posture.
VP of Engineering, Later