Skip to content
Course content preview

Evaluating retrieval quality

Because an answer that was never retrieved cannot be generated. Recall@k, the share of the answering documents that a retriever puts in its top k results, is the ceiling on everything downstream, and no prompt, model or reranker lifts it. So when a RAG answer comes out wrong, the first thing to check is whether the right document reached the context at all. That is a retrieval measurement, and it is separate from anything the generation model did with what it was given.

The course has made two notable claims so far:

  • Lesson 05 said hybrid search beats either half of it, and

  • Lesson 09 said a smaller chunk retrieves better than a large one.

However, neither claim has been tested, which is why this lesson builds the harness that tests them, and reports what it finds even where that disagrees with the lesson making the claim.

The corpus, the analyzer and the two indexes are lesson 05's, unchanged: ten help-centre articles for a fictional online store, a BM25 full-text index and an HNSW vector index. This lesson adds no retrieval feature, and measures the ones you already have.

An in-memory instance in its own terminal is authenticated and throwaway:

surreal start --user root --pass secret

surreal import --endpoint http://localhost:8000 \
  --user root --pass secret --ns ai --db eval schema.surql
surreal import --endpoint http://localhost:8000 \
  --user root --pass secret --ns ai --db eval seed.surql

Three files go in: schema.surql for the corpus, the golden set, three retrievers and the metrics, seed.surql for ten articles and eight labelled questions, and queries.surql for the evaluation.

A golden set is a list of questions with the documents that answer them, labelled by a human. Most teams keep it in a spreadsheet or a JSON file. Keep it in the database instead and the labels can be typed:

DEFINE TABLE OVERWRITE question SCHEMAFULL;

DEFINE FIELD OVERWRITE text      ON question TYPE string;
DEFINE FIELD OVERWRITE embedding ON question TYPE array<float>
    ASSERT $value.len() = 4;
DEFINE FIELD OVERWRITE relevant  ON question TYPE array<record<article>>
    ASSERT $value.len() > 0;

array<record<article>> means every label has to point at an article rather than at some other table, which an evaluation set kept in a spreadsheet cannot check at all. The type does not check that the article still exists, though: delete one and the label sits there pointing at nothing. Adding REFERENCE ON DELETE REJECT to the field closes that gap, refusing to delete an article any question still depends on, and ON DELETE UNSET is the other option, which drops the deleted article from the label instead.

The type still allows an empty array, which is what the ASSERT is for. A question nobody labelled scores zero for every retriever and quietly drags every average down.

The eight questions are phrased the way a customer would phrase them, not the way the articles are written:

CREATE question:signin SET
    text = "I can't get into my account",
    embedding = [0.95, 0.0, 0.0, 0.1],
    relevant = [article:password_reset, article:account_locked, article:enable_2fa];

CREATE question:rate_limit SET
    text = "429 too many requests",
    embedding = [0.25, 0.25, 0.25, 0.25],
    relevant = [article:api_rate_limits];

Two of those numbers are the whole experiment. The sign-in question shares no vocabulary with any of its three answers, so it can only be found semantically. The 429 question is a bare error code, whose embedding is a vague direction rather than a topic, so its vector is near-uniform on purpose. Lesson 05 makes the same modelling choice for the same reason.

Reading the set back gives you the review artefact, which is the other reason to keep it here (three of the eight records):

SELECT id, text, relevant.title AS answers FROM question ORDER BY id;
Output
[
    { id: question:money_back, text: 'how do I get my money back',
      answers: ['Request a refund', 'Cancel or pause a subscription'] },
    { id: question:rate_limit, text: '429 too many requests',
      answers: ['API rate limits and 429 errors'] },
    { id: question:signin,     text: "I can't get into my account",
      answers: ['Reset a forgotten password', 'Why is my account locked?', 'Set up two-factor authentication'] }
]

A metric needs nothing from a retriever except record IDs, best first. Reduce every candidate to that shape and they become comparable:

-- Semantic only, over the HNSW index.
DEFINE FUNCTION OVERWRITE fn::retrieve_vec($qvec: array<float>, $k: int) -> array {
    RETURN (SELECT id, vector::distance::knn() AS d FROM article
        WHERE embedding <|10, 40|> $qvec ORDER BY d ASC LIMIT $k).id;
};

-- Lexical only. '@1,OR@' matches any term and exposes BM25 as search::score(1).
DEFINE FUNCTION OVERWRITE fn::retrieve_bm25($q: string, $k: int) -> array {
    RETURN (SELECT id, search::score(1) AS s FROM article
        WHERE content @1,OR@ $q ORDER BY s DESC LIMIT $k).id;
};

-- Both, fused on rank with the Reciprocal Rank Fusion from lesson 05.
DEFINE FUNCTION OVERWRITE fn::retrieve_hybrid(
    $q: string, $qvec: array<float>, $k: int
) -> array {
    LET $bm25 = SELECT id, search::score(1) AS s FROM article
        WHERE content @1,OR@ $q ORDER BY s DESC LIMIT 10;
    LET $vec = SELECT id, vector::distance::knn() AS d FROM article
        WHERE embedding <|10, 40|> $qvec ORDER BY d ASC;
    RETURN search::rrf([$bm25, $vec], $k, 60).id;
};

The KNN operator's K has to be a literal, so fn::retrieve_vec asks for the whole corpus and lets the caller's $k do the cutting. On a real corpus you would set that literal to your largest candidate pool.

Here are the three on the 429 question:

LET $qn = question:rate_limit;
RETURN fn::retrieve_bm25($qn.text, 3);
RETURN fn::retrieve_vec($qn.embedding, 3);
RETURN fn::retrieve_hybrid($qn.text, $qn.embedding, 3);
Output
[article:api_rate_limits, article:account_locked, article:refund_request]
[article:return_item, article:cancel_subscription, article:enable_2fa]
[article:refund_request, article:account_locked, article:api_rate_limits]

The labelled answer is article:api_rate_limits. BM25 puts it first, the vector retriever misses it entirely, and hybrid recovers it in third place.

There are three, and they fit in a screen. recall@k asks how much of the answer you found:

-- recall@k: of the articles that answer this question, how many are in the top k?
DEFINE FUNCTION OVERWRITE fn::recall_at($retrieved: array, $relevant: array, $k: int) -> float {
    LET $top = $retrieved[0..math::min([$k, $retrieved.len()])];
    RETURN <float> $top.intersect($relevant).len() / $relevant.len();
};

precision@k asks what fraction of the context you spent was worth spending:

-- precision@k: of the articles sent to the model, how many belonged there?
-- Every irrelevant one is context budget spent on nothing.
DEFINE FUNCTION OVERWRITE fn::precision_at($retrieved: array, $relevant: array, $k: int) -> float {
    LET $top = $retrieved[0..math::min([$k, $retrieved.len()])];
    RETURN IF $top.len() = 0 { 0f } ELSE { <float> $top.intersect($relevant).len() / $top.len() };
};

Reciprocal rank asks where the first correct hit landed, and averaged over a set it is the mean reciprocal rank, or MRR:

-- Reciprocal rank: 1 / the position of the first relevant hit, or 0 if there is none.
DEFINE FUNCTION OVERWRITE fn::rr($retrieved: array, $relevant: array) -> float {
    LET $first = $retrieved.map(|$id| $relevant CONTAINS $id).find_index(true);
    RETURN IF $first = NONE { 0f } ELSE { 1f / ($first + 1) };
};

The math::min([$k, $retrieved.len()]) is there because a range index past the end of an array evaluates to NONE rather than to the short array, so a retriever that returned fewer than k results would silently score zero without that clamp. This is the kind of bug an eval harness must avoid.

This is what the semantic retriever returns for the 429 question:

LET $got = fn::retrieve_vec($qn.embedding, 3);
RETURN { recall_at_3: fn::recall_at($got, $qn.relevant, 3),
    precision_at_3: fn::precision_at($got, $qn.relevant, 3),
    rr: fn::rr($got, $qn.relevant) };
Output
{ recall_at_3: 0f, precision_at_3: 0f, rr: 0f }

fn::evaluate($k) runs all eight questions against all three retrievers and returns the means:

RETURN [1, 3, 5].map(|$k| fn::evaluate($k));
Output
[
    { k: 1, bm25: 0.438f, vector: 0.667f, hybrid: 0.479f },
    { k: 3, bm25: 0.771f, vector: 0.875f, hybrid: 0.833f },
    { k: 5, bm25: 0.833f, vector: 0.875f, hybrid: 0.958f }
]

That table does not say what lesson 05 led you to expect, so let's take a look at why that is the case.

At k = 3, the hybrid retriever loses to vector search alone, 0.833 against 0.875. At k = 5 it wins, 0.958 against 0.875. Fusion needs room: with only three slots, the lexical half spends one of them on a candidate the semantic half would not have chosen, and on this corpus that trade is a loss. Raise k and the same trade becomes a gain.

MRR tells a consistent story at k = 5:

Output
[{ mrr_bm25: 0.646f, mrr_vec: 0.875f, mrr_hybrid: 0.76f }]

Hybrid finds more of the answers than vector search and puts the first one lower down. If your prompt takes five documents, that is fine. If only the top one is ever used, it is not. The metric you choose has to match the shape of what you build.

The same run gives one record per question:

SELECT text, relevant.len() AS labels,
    math::round(fn::recall_at(fn::retrieve_bm25(text, 3), relevant, 3) * 1000) / 1000 AS bm25,
    math::round(fn::recall_at(fn::retrieve_vec(embedding, 3), relevant, 3) * 1000) / 1000 AS vec,
    math::round(fn::recall_at(fn::retrieve_hybrid(text, embedding, 3), relevant, 3) * 1000) / 1000 AS hybrid
FROM question;
questionlabelsbm25vectorhybrid
my credit card was declined11.01.01.0
how do I get my money back20.01.00.0
where is my parcel11.01.01.0
429 too many requests11.00.01.0
I want to send this item back20.51.01.0
I can't get into my account30.6671.00.667
stop charging me every month11.01.01.0
how do I verify a webhook payload11.01.01.0

The 0.042 the mean moved between vector and hybrid is two questions cancelling out. Hybrid fixed the 429 question completely, from 0 to 1.0, which is the case lesson 05 was built for. It also broke "how do I get my money back" completely, from 1.0 to 0.

A mean would have hidden both. Always keep the per-question records: the aggregate tells you whether the change is an improvement, and the individual records tell you what to fix.

LET $bad = question:money_back;
RETURN { question: $bad.text, labelled: $bad.relevant.title,
    vector: fn::retrieve_vec($bad.embedding, 3).title,
    hybrid: fn::retrieve_hybrid($bad.text, $bad.embedding, 3).title };
Output
{
    question: 'how do I get my money back',
    labelled: ['Request a refund', 'Cancel or pause a subscription'],
    vector:   ['Request a refund', 'Cancel or pause a subscription', 'Update your payment method'],
    hybrid:   ['Return an item', 'Track your order', 'API rate limits and 429 errors']
}

Vector search got it exactly right. Hybrid returned three articles, none of them labelled. The lexical list is where to look:

SELECT title, math::round(search::score(1) * 100) / 100 AS bm25
FROM article WHERE content @1,OR@ "how do I get my money back" ORDER BY bm25 DESC;
Output
[
    { title: 'Return an item',                 bm25: 2.46f },
    { title: 'Track your order',               bm25: 1.27f },
    { title: 'API rate limits and 429 errors', bm25: 1.15f }
]

There it is. "Return an item" contains "send something back", and @1,OR@ matches any term, so a question about money matches an article about parcels on one word. Three weak matches come back, all of them wrong and all of them ranked.

RRF fuses on rank and throws the scores away, which is what makes it robust when both lists are decent. The cost is that it cannot tell corroboration from coincidence: a rank-1 candidate in a list of pure noise carries the same weight as a rank-1 candidate in a good list. Two of these noise articles also appear low in the vector list, so they collect points from both sides and outvote the right answers.

The hypothesis: BM25 candidates that score 2.46 on a six-word question are noise, and fusing them pushes better candidates out of the top k. So require a lexical candidate to clear a floor before it is fused at all:

-- Both, fused on rank, with a floor on the lexical list: a BM25 candidate has to
-- score at least $floor to be fused at all.
DEFINE FUNCTION OVERWRITE fn::retrieve_hybrid_floor(
    $q: string, $qvec: array<float>, $k: int, $floor: float
) -> array {
    LET $bm25 = SELECT id, search::score(1) AS s FROM article
        WHERE content @1,OR@ $q ORDER BY s DESC LIMIT 10;
    LET $vec = SELECT id, vector::distance::knn() AS d FROM article
        WHERE embedding <|10, 40|> $qvec ORDER BY d ASC;
    RETURN search::rrf([$bm25.filter(|$r| $r.s >= $floor), $vec], $k, 60).id;
};

Now it is a measurement rather than an opinion:

RETURN [0f, 1.5f, 2.5f, 4f].map(|$floor| fn::evaluate_floor(3, $floor));
Output
[
    { k: 3, floor: 0f,   recall: 0.833f },
    { k: 3, floor: 1.5f, recall: 1f },
    { k: 3, floor: 2.5f, recall: 1f },
    { k: 3, floor: 4f,   recall: 1f }
]

Recall@3 goes from 0.833 to 1.0, beating both halves and the unfloored fusion. At k = 5 it goes from 0.958 to 1.0.

The uncomfortable part is also the most useful thing in this lesson. Every floor from 1.5 to 4.0 scores identically. Eight questions cannot tell them apart, so picking 2.5 because it is in the middle is a coin flip with a decimal point on it. And the floor was chosen by looking at the failure it fixes, on the same set used to score it, which is how you overfit an evaluation set. The number to report is from questions the change has not seen.

An evaluation you run once tells you where you stand today. An evaluation you keep tells you when something has got worse, because each run can be compared against the ones before it:

DEFINE TABLE OVERWRITE eval_run SCHEMAFULL;
DEFINE FIELD OVERWRITE at        ON eval_run TYPE datetime DEFAULT time::now();
DEFINE FIELD OVERWRITE retriever ON eval_run TYPE string;
DEFINE FIELD OVERWRITE k         ON eval_run TYPE int;
DEFINE FIELD OVERWRITE recall    ON eval_run TYPE float;
DEFINE FIELD OVERWRITE questions ON eval_run TYPE int;
LET $n = (SELECT count() FROM question GROUP ALL)[0].count;
CREATE eval_run SET retriever = 'hybrid+floor', k = 3,
    recall = fn::evaluate_floor(3, 1.5f).recall, questions = $n;
SELECT retriever, k, recall, questions FROM eval_run ORDER BY recall DESC;
Output
[
    { retriever: 'hybrid+floor', k: 3, recall: 1f,     questions: 8 },
    { retriever: 'vector',       k: 3, recall: 0.875f, questions: 8 },
    { retriever: 'hybrid',       k: 3, recall: 0.833f, questions: 8 },
    { retriever: 'bm25',         k: 3, recall: 0.771f, questions: 8 }
]

The history lives in the same database as the corpus it measures, so a run is one CREATE and a trend is one SELECT. If you run queries.surql twice you get two sets of records. The questions field is there so a number from a 40-question set is never compared with a number from an 8-question one.

From here it can run on its own. Wrap the run in a function and call it whenever documents are ingested, so a drop in the numbers appears next to the data that caused it. Nothing has to leave the database, and what is measured cannot drift away from what does the measuring.

It needs to be bigger than eight, and the reason is resolution rather than any particular threshold. Recall@k is not binary once a question has more than one relevant document - the run in step 6 scored 0.5 and 0.667 on questions labelled with two and three documents - so the step a single question contributes to the mean is 1 / (labels × question_count), not a flat 1 / question_count. A retrieval change on "I can't get into my account" (3 labels) moves the eight-question mean by 1 / (3 × 8), about 4.17 percentage points; the same kind of change on a single-label question like "429 too many requests" moves it the full 1/8, or 12.5 points. Step 8 is what the coarse end of that range looks like: four different floors scored identically because the gaps between them were smaller than even the single-label step. Adding questions shrinks that coarse step - still roughly 1 / question_count - and questions with more labels sharpen the mean's resolution further on top of that, so the set does not cross a line at some number, it just gets finer.

The general advice in statistics that a sample should be at least 30 comes from a different question, which is roughly how many samples make the distribution of a mean close enough to normal to put a confidence interval around it. It is a convention rather than a boundary, and it is answering "how certain am I of this number" where the paragraph above answers "how small a difference can I see at all". Both happen to point at a few dozen questions.

Some working numbers:

  • 30 to 50 questions is the smallest set where a mean starts to mean anything, and one question still moves it 2 to 3 points.

  • 100 to 200 if you want to slice by category or question type and read the slices.

  • Every question that ever failed in production belongs in the set permanently. That is the part that stops the same regression shipping twice.

Where questions come from matters more than how many there are. Take them from support tickets, search logs and the questions users actually asked. Questions written by the person who wrote the corpus test the corpus against itself, and they always score better than reality.

Label what actually answers the question, not everything on the topic. Every extra label lowers recall for a retriever that was right, and the metric then rewards a retriever that returns broad, vague context.

Four things sit outside what it measures:

Whether these numbers transfer. The corpus is ten documents with hand-written 4-dimensional vectors. That makes semantic search unrealistically strong: the question vectors were written to point at the topics the articles cover, with none of the noise a real model produces. Vector search scoring 0.875 here says nothing about your corpus. Run this harness on your data, with your model, before you rely on any of these figures.

Whether the answer was right. Recall@k is a ceiling, not a guarantee. A model handed the correct document can still contradict it. Faithfulness and answer quality need their own evaluation, with the generated answer in the loop.

Whether it is fast enough, or cheap enough. Neither appears in any of these metrics. The floored hybrid retriever runs two index scans and a fusion for every query.

Whether the labels are right. The golden set is a human artefact and it goes stale as the corpus changes. While array<record<article>> stops a label pointing at a non-article, nothing stops a label pointing at an article that has since been rewritten to say something else.

  1. 1.

    Replace the corpus with your documents, and DIMENSION 4 with your model's dimensionality.

  2. 2.

    Write 30 questions from your support log, and label them by hand.

  3. 3.

    Keep fn::retrieve_*, the three metrics and fn::evaluate as they are. They only assume a ranked list of record IDs.

  4. 4.

    Put the run in CI, on a fixed k that matches how many documents your prompt actually takes.

Then use it on the decision from lesson 09. Chunk your corpus two ways into two tables, point the retriever at each, and let recall@k choose. That gives you a number for the choice, which is more than any chunking guide can offer.

THE PLATFORM

Everything an application and its agents know. Five surfaces, one engine.

IN PRODUCTION

Trusted at scale. Samsung, Nvidia, Verizon, Tencent, and Walmart run on SurrealDB.

14,000+

Developers building on SurrealDB Cloud

4M+

Developers building on SurrealDB worldwide

FROM THE TEAMS

SurrealDB gives us a foundation where we can unify semantic search, knowledge graphs, and AI-driven decision making without stitching together multiple systems. Collapsing responsibility into SurrealDB has become our default engineering posture.
Justin Foley

VP of Engineering, Later

SurrealDB

The context and memory layer for AI agents

Database. Graphs, vectors, documents and relational data in one engine, in a single ACID transaction.
Agent Memory. Connects and retrieves context wherever your data lives, every fact carrying its source.
Cloud. Fully managed, in the cloud provider and region you choose.

Explore with AI

Copyright © 2026 SurrealDB Ltd. Registered in England and Wales. Company no. 13615201

Registered address: 3rd Floor 1 Ashley Road, Altrincham, Cheshire, WA14 2DT, United Kingdom

Trading address: Huckletree Oxford Circus, 213 Oxford Street, London, W1D 2LG, United Kingdom