---
title: "Evaluating retrieval quality | SurrealDB University"
description: "Build a golden set as a table, compute recall@k, precision@k and MRR in SurrealQL, and measure three retrievers against each other without leaving the database."
url: https://surrealdb.com/learn/ai/evaluating-retrieval
---

![Course content preview](https://surrealdb.com/assets/static/course-ai.LEsq_J_G.avif)

Course chapters

[Back to courses](https://surrealdb.com/learn) [SurrealDB for AI Engineers](https://surrealdb.com/learn/ai) [AI foundations](https://surrealdb.com/learn/ai/ai-foundations) [Vector embeddings and search](https://surrealdb.com/learn/ai/vector-embeddings) [Full-text search and BM25](https://surrealdb.com/learn/ai/fulltext-search-bm25) [Building a RAG knowledge base](https://surrealdb.com/learn/ai/rag-knowledge-base) [Hybrid search and reranking](https://surrealdb.com/learn/ai/hybrid-search-reranking) [Building an agent memory store](https://surrealdb.com/learn/ai/agent-memory-store) [Text-to-SurQL: agentic prompt engineering](https://surrealdb.com/learn/ai/text-to-surrealql) [Many sources and agents, one context layer](https://surrealdb.com/learn/ai/multi-source-context-layer) [Chunking strategies](https://surrealdb.com/learn/ai/chunking-strategies) [Evaluating retrieval quality](https://surrealdb.com/learn/ai/evaluating-retrieval) [Graph RAG beyond one hop](https://surrealdb.com/learn/ai/graph-rag-multi-hop) Certificate Pending

# Evaluating retrieval quality

## Why measure retrieval separately from the answer?

Because an answer that was never retrieved cannot be generated. Recall@k, the share of the answering documents that a retriever puts in its top k results, is the ceiling on everything downstream, and no prompt, model or reranker lifts it. So when a RAG answer comes out wrong, the first thing to check is whether the right document reached the context at all. That is a retrieval measurement, and it is separate from anything the generation model did with what it was given.

The course has made two notable claims so far:

- [Lesson 05](https://surrealdb.com/learn/ai/hybrid-search-reranking) said hybrid search beats either half of it, and
- [Lesson 09](https://surrealdb.com/learn/ai/chunking-strategies) said a smaller chunk retrieves better than a large one.

However, neither claim has been tested, which is why this lesson builds the harness that tests them, and reports what it finds even where that disagrees with the lesson making the claim.

## Step 1 - start a server

The corpus, the analyzer and the two indexes are [lesson 05](https://surrealdb.com/learn/ai/hybrid-search-reranking)'s, unchanged: ten help-centre articles for a fictional online store, a BM25 full-text index and an HNSW vector index. This lesson adds no retrieval feature, and measures the ones you already have.

An in-memory instance in its own terminal is authenticated and throwaway:

```bash
surreal start --user root --pass secret

surreal import --endpoint http://localhost:8000 \
  --user root --pass secret --ns ai --db eval schema.surql
surreal import --endpoint http://localhost:8000 \
  --user root --pass secret --ns ai --db eval seed.surql
```

Three files go in: [`schema.surql`](https://gist.github.com/martinschaer/ad587291d66dcc836c32a14911bc9f55) for the corpus, the golden set, three retrievers and the metrics, [`seed.surql`](https://gist.github.com/martinschaer/3624f32d9c42317d147d8c6f4d8f5b69) for ten articles and eight labelled questions, and [`queries.surql`](https://gist.github.com/martinschaer/ee90369014d0c11ec7c635cc1f9b466c) for the evaluation.

## Step 2 - the golden set is a table

A golden set is a list of questions with the documents that answer them, labelled by a human. Most teams keep it in a spreadsheet or a JSON file. Keep it in the database instead and the labels can be typed:

```surql
DEFINE TABLE OVERWRITE question SCHEMAFULL;

DEFINE FIELD OVERWRITE text      ON question TYPE string;
DEFINE FIELD OVERWRITE embedding ON question TYPE array<float>
    ASSERT $value.len() = 4;
DEFINE FIELD OVERWRITE relevant  ON question TYPE array<record<article>>
    ASSERT $value.len() > 0;
```

`array<record<article>>` means every label has to point at an article rather than at some other table, which an evaluation set kept in a spreadsheet cannot check at all. The type does not check that the article still exists, though: delete one and the label sits there pointing at nothing. Adding `REFERENCE ON DELETE REJECT` to the field closes that gap, refusing to delete an article any question still depends on, and `ON DELETE UNSET` is the other option, which drops the deleted article from the label instead.

The type still allows an empty array, which is what the `ASSERT` is for. A question nobody labelled scores zero for every retriever and quietly drags every average down.

The eight questions are phrased the way a customer would phrase them, not the way the articles are written:

```surql
CREATE question:signin SET
    text = "I can't get into my account",
    embedding = [0.95, 0.0, 0.0, 0.1],
    relevant = [article:password_reset, article:account_locked, article:enable_2fa];

CREATE question:rate_limit SET
    text = "429 too many requests",
    embedding = [0.25, 0.25, 0.25, 0.25],
    relevant = [article:api_rate_limits];
```

Two of those numbers are the whole experiment. The sign-in question shares no vocabulary with any of its three answers, so it can only be found semantically. The `429` question is a bare error code, whose embedding is a vague direction rather than a topic, so its vector is near-uniform on purpose. [Lesson 05](https://surrealdb.com/learn/ai/hybrid-search-reranking) makes the same modelling choice for the same reason.

Reading the set back gives you the review artefact, which is the other reason to keep it here (three of the eight records):

```surql
SELECT id, text, relevant.title AS answers FROM question ORDER BY id;
```

Output

```surql
[
    { id: question:money_back, text: 'how do I get my money back',
      answers: ['Request a refund', 'Cancel or pause a subscription'] },
    { id: question:rate_limit, text: '429 too many requests',
      answers: ['API rate limits and 429 errors'] },
    { id: question:signin,     text: "I can't get into my account",
      answers: ['Reset a forgotten password', 'Why is my account locked?', 'Set up two-factor authentication'] }
]
```

## Step 3 - one shape for every retriever

A metric needs nothing from a retriever except record IDs, best first. Reduce every candidate to that shape and they become comparable:

```surql
-- Semantic only, over the HNSW index.
DEFINE FUNCTION OVERWRITE fn::retrieve_vec($qvec: array<float>, $k: int) -> array {
    RETURN (SELECT id, vector::distance::knn() AS d FROM article
        WHERE embedding <|10, 40|> $qvec ORDER BY d ASC LIMIT $k).id;
};

-- Lexical only. '@1,OR@' matches any term and exposes BM25 as search::score(1).
DEFINE FUNCTION OVERWRITE fn::retrieve_bm25($q: string, $k: int) -> array {
    RETURN (SELECT id, search::score(1) AS s FROM article
        WHERE content @1,OR@ $q ORDER BY s DESC LIMIT $k).id;
};

-- Both, fused on rank with the Reciprocal Rank Fusion from lesson 05.
DEFINE FUNCTION OVERWRITE fn::retrieve_hybrid(
    $q: string, $qvec: array<float>, $k: int
) -> array {
    LET $bm25 = SELECT id, search::score(1) AS s FROM article
        WHERE content @1,OR@ $q ORDER BY s DESC LIMIT 10;
    LET $vec = SELECT id, vector::distance::knn() AS d FROM article
        WHERE embedding <|10, 40|> $qvec ORDER BY d ASC;
    RETURN search::rrf([$bm25, $vec], $k, 60).id;
};
```

The KNN operator's `K` has to be a literal, so `fn::retrieve_vec` asks for the whole corpus and lets the caller's `$k` do the cutting. On a real corpus you would set that literal to your largest candidate pool.

Here are the three on the `429` question:

```surql
LET $qn = question:rate_limit;
RETURN fn::retrieve_bm25($qn.text, 3);
RETURN fn::retrieve_vec($qn.embedding, 3);
RETURN fn::retrieve_hybrid($qn.text, $qn.embedding, 3);
```

Output

```surql
[article:api_rate_limits, article:account_locked, article:refund_request]
[article:return_item, article:cancel_subscription, article:enable_2fa]
[article:refund_request, article:account_locked, article:api_rate_limits]
```

The labelled answer is `article:api_rate_limits`. BM25 puts it first, the vector retriever misses it entirely, and hybrid recovers it in third place.

## Step 4 - the metrics are functions

There are three, and they fit in a screen. **recall@k** asks how much of the answer you found:

```surql
-- recall@k: of the articles that answer this question, how many are in the top k?
DEFINE FUNCTION OVERWRITE fn::recall_at($retrieved: array, $relevant: array, $k: int) -> float {
    LET $top = $retrieved[0..math::min([$k, $retrieved.len()])];
    RETURN <float> $top.intersect($relevant).len() / $relevant.len();
};
```

**precision@k** asks what fraction of the context you spent was worth spending:

```surql
-- precision@k: of the articles sent to the model, how many belonged there?
-- Every irrelevant one is context budget spent on nothing.
DEFINE FUNCTION OVERWRITE fn::precision_at($retrieved: array, $relevant: array, $k: int) -> float {
    LET $top = $retrieved[0..math::min([$k, $retrieved.len()])];
    RETURN IF $top.len() = 0 { 0f } ELSE { <float> $top.intersect($relevant).len() / $top.len() };
};
```

**Reciprocal rank** asks where the first correct hit landed, and averaged over a set it is the **mean reciprocal rank**, or MRR:

```surql
-- Reciprocal rank: 1 / the position of the first relevant hit, or 0 if there is none.
DEFINE FUNCTION OVERWRITE fn::rr($retrieved: array, $relevant: array) -> float {
    LET $first = $retrieved.map(|$id| $relevant CONTAINS $id).find_index(true);
    RETURN IF $first = NONE { 0f } ELSE { 1f / ($first + 1) };
};
```

The `math::min([$k, $retrieved.len()])` is there because a range index past the end of an array evaluates to `NONE` rather than to the short array, so a retriever that returned fewer than `k` results would silently score zero without that clamp. This is the kind of bug an eval harness must avoid.

This is what the semantic retriever returns for the `429` question:

```surql
LET $got = fn::retrieve_vec($qn.embedding, 3);
RETURN { recall_at_3: fn::recall_at($got, $qn.relevant, 3),
    precision_at_3: fn::precision_at($got, $qn.relevant, 3),
    rr: fn::rr($got, $qn.relevant) };
```

Output

```surql
{ recall_at_3: 0f, precision_at_3: 0f, rr: 0f }
```

## Step 5 - run the set

`fn::evaluate($k)` runs all eight questions against all three retrievers and returns the means:

```surql
RETURN [1, 3, 5].map(|$k| fn::evaluate($k));
```

Output

```surql
[
    { k: 1, bm25: 0.438f, vector: 0.667f, hybrid: 0.479f },
    { k: 3, bm25: 0.771f, vector: 0.875f, hybrid: 0.833f },
    { k: 5, bm25: 0.833f, vector: 0.875f, hybrid: 0.958f }
]
```

That table does not say what [lesson 05](https://surrealdb.com/learn/ai/hybrid-search-reranking) led you to expect, so let's take a look at why that is the case.

At k = 3, the hybrid retriever **loses to vector search alone**, 0.833 against 0.875. At k = 5 it wins, 0.958 against 0.875. Fusion needs room: with only three slots, the lexical half spends one of them on a candidate the semantic half would not have chosen, and on this corpus that trade is a loss. Raise k and the same trade becomes a gain.

MRR tells a consistent story at k = 5:

Output

```surql
[{ mrr_bm25: 0.646f, mrr_vec: 0.875f, mrr_hybrid: 0.76f }]
```

Hybrid finds more of the answers than vector search and puts the first one lower down. If your prompt takes five documents, that is fine. If only the top one is ever used, it is not. The metric you choose has to match the shape of what you build.

## Step 6 - the mean hides what you need to see

The same run gives one record per question:

```surql
SELECT text, relevant.len() AS labels,
    math::round(fn::recall_at(fn::retrieve_bm25(text, 3), relevant, 3) * 1000) / 1000 AS bm25,
    math::round(fn::recall_at(fn::retrieve_vec(embedding, 3), relevant, 3) * 1000) / 1000 AS vec,
    math::round(fn::recall_at(fn::retrieve_hybrid(text, embedding, 3), relevant, 3) * 1000) / 1000 AS hybrid
FROM question;
```

| question | labels | bm25 | vector | hybrid |
| --- | --- | --- | --- | --- |
| my credit card was declined | 1 | 1.0 | 1.0 | 1.0 |
| how do I get my money back | 2 | 0.0 | **1.0** | **0.0** |
| where is my parcel | 1 | 1.0 | 1.0 | 1.0 |
| 429 too many requests | 1 | 1.0 | **0.0** | **1.0** |
| I want to send this item back | 2 | 0.5 | 1.0 | 1.0 |
| I can't get into my account | 3 | 0.667 | 1.0 | 0.667 |
| stop charging me every month | 1 | 1.0 | 1.0 | 1.0 |
| how do I verify a webhook payload | 1 | 1.0 | 1.0 | 1.0 |

The 0.042 the mean moved between vector and hybrid is two questions cancelling out. Hybrid fixed the `429` question completely, from 0 to 1.0, which is the case [lesson 05](https://surrealdb.com/learn/ai/hybrid-search-reranking) was built for. It also broke "how do I get my money back" completely, from 1.0 to 0.

A mean would have hidden both. Always keep the per-question records: the aggregate tells you whether the change is an improvement, and the individual records tell you what to fix.

## Step 7 - diagnose the one that broke

```surql
LET $bad = question:money_back;
RETURN { question: $bad.text, labelled: $bad.relevant.title,
    vector: fn::retrieve_vec($bad.embedding, 3).title,
    hybrid: fn::retrieve_hybrid($bad.text, $bad.embedding, 3).title };
```

Output

```surql
{
    question: 'how do I get my money back',
    labelled: ['Request a refund', 'Cancel or pause a subscription'],
    vector:   ['Request a refund', 'Cancel or pause a subscription', 'Update your payment method'],
    hybrid:   ['Return an item', 'Track your order', 'API rate limits and 429 errors']
}
```

Vector search got it exactly right. Hybrid returned three articles, none of them labelled. The lexical list is where to look:

```surql
SELECT title, math::round(search::score(1) * 100) / 100 AS bm25
FROM article WHERE content @1,OR@ "how do I get my money back" ORDER BY bm25 DESC;
```

Output

```surql
[
    { title: 'Return an item',                 bm25: 2.46f },
    { title: 'Track your order',               bm25: 1.27f },
    { title: 'API rate limits and 429 errors', bm25: 1.15f }
]
```

There it is. "Return an item" contains "send something **back**", and `@1,OR@` matches any term, so a question about money matches an article about parcels on one word. Three weak matches come back, all of them wrong and all of them ranked.

RRF fuses on rank and throws the scores away, which is what makes it robust when both lists are decent. The cost is that it cannot tell corroboration from coincidence: a rank-1 candidate in a list of pure noise carries the same weight as a rank-1 candidate in a good list. Two of these noise articles also appear low in the vector list, so they collect points from both sides and outvote the right answers.

## Step 8 - measure the fix

The hypothesis: BM25 candidates that score 2.46 on a six-word question are noise, and fusing them pushes better candidates out of the top k. So require a lexical candidate to clear a floor before it is fused at all:

```surql
-- Both, fused on rank, with a floor on the lexical list: a BM25 candidate has to
-- score at least $floor to be fused at all.
DEFINE FUNCTION OVERWRITE fn::retrieve_hybrid_floor(
    $q: string, $qvec: array<float>, $k: int, $floor: float
) -> array {
    LET $bm25 = SELECT id, search::score(1) AS s FROM article
        WHERE content @1,OR@ $q ORDER BY s DESC LIMIT 10;
    LET $vec = SELECT id, vector::distance::knn() AS d FROM article
        WHERE embedding <|10, 40|> $qvec ORDER BY d ASC;
    RETURN search::rrf([$bm25.filter(|$r| $r.s >= $floor), $vec], $k, 60).id;
};
```

Now it is a measurement rather than an opinion:

```surql
RETURN [0f, 1.5f, 2.5f, 4f].map(|$floor| fn::evaluate_floor(3, $floor));
```

Output

```surql
[
    { k: 3, floor: 0f,   recall: 0.833f },
    { k: 3, floor: 1.5f, recall: 1f },
    { k: 3, floor: 2.5f, recall: 1f },
    { k: 3, floor: 4f,   recall: 1f }
]
```

Recall@3 goes from 0.833 to 1.0, beating both halves and the unfloored fusion. At k = 5 it goes from 0.958 to 1.0.

The uncomfortable part is also the most useful thing in this lesson. **Every floor from 1.5 to 4.0 scores identically.** Eight questions cannot tell them apart, so picking 2.5 because it is in the middle is a coin flip with a decimal point on it. And the floor was chosen by looking at the failure it fixes, on the same set used to score it, which is how you overfit an evaluation set. The number to report is from questions the change has not seen.

## Step 9 - keep the run

An evaluation you run once tells you where you stand today. An evaluation you keep tells you when something has got worse, because each run can be compared against the ones before it:

```surql
DEFINE TABLE OVERWRITE eval_run SCHEMAFULL;
DEFINE FIELD OVERWRITE at        ON eval_run TYPE datetime DEFAULT time::now();
DEFINE FIELD OVERWRITE retriever ON eval_run TYPE string;
DEFINE FIELD OVERWRITE k         ON eval_run TYPE int;
DEFINE FIELD OVERWRITE recall    ON eval_run TYPE float;
DEFINE FIELD OVERWRITE questions ON eval_run TYPE int;
```

```surql
LET $n = (SELECT count() FROM question GROUP ALL)[0].count;
CREATE eval_run SET retriever = 'hybrid+floor', k = 3,
    recall = fn::evaluate_floor(3, 1.5f).recall, questions = $n;
SELECT retriever, k, recall, questions FROM eval_run ORDER BY recall DESC;
```

Output

```surql
[
    { retriever: 'hybrid+floor', k: 3, recall: 1f,     questions: 8 },
    { retriever: 'vector',       k: 3, recall: 0.875f, questions: 8 },
    { retriever: 'hybrid',       k: 3, recall: 0.833f, questions: 8 },
    { retriever: 'bm25',         k: 3, recall: 0.771f, questions: 8 }
]
```

The history lives in the same database as the corpus it measures, so a run is one `CREATE` and a trend is one `SELECT`. If you run [`queries.surql`](https://gist.github.com/martinschaer/ee90369014d0c11ec7c635cc1f9b466c) twice you get two sets of records. The `questions` field is there so a number from a 40-question set is never compared with a number from an 8-question one.

From here it can run on its own. Wrap the run in a function and call it whenever documents are ingested, so a drop in the numbers appears next to the data that caused it. Nothing has to leave the database, and what is measured cannot drift away from what does the measuring.

## How big does a golden set need to be?

It needs to be bigger than eight, and the reason is resolution rather than any particular threshold. Recall@k is not binary once a question has more than one relevant document - the run in step 6 scored 0.5 and 0.667 on questions labelled with two and three documents - so the step a single question contributes to the mean is 1 / (labels × question_count), not a flat 1 / question_count. A retrieval change on "I can't get into my account" (3 labels) moves the eight-question mean by 1 / (3 × 8), about 4.17 percentage points; the same kind of change on a single-label question like "429 too many requests" moves it the full 1/8, or 12.5 points. Step 8 is what the coarse end of that range looks like: four different floors scored identically because the gaps between them were smaller than even the single-label step. Adding questions shrinks that coarse step - still roughly 1 / question_count - and questions with more labels sharpen the mean's resolution further on top of that, so the set does not cross a line at some number, it just gets finer.

The general advice in statistics that a sample should be at least 30 comes from a different question, which is roughly how many samples make the distribution of a mean close enough to normal to put a confidence interval around it. It is a convention rather than a boundary, and it is answering "how certain am I of this number" where the paragraph above answers "how small a difference can I see at all". Both happen to point at a few dozen questions.

Some working numbers:

- **30 to 50 questions** is the smallest set where a mean starts to mean anything, and one question still moves it 2 to 3 points.
- **100 to 200** if you want to slice by category or question type and read the slices.
- **Every question that ever failed in production** belongs in the set permanently. That is the part that stops the same regression shipping twice.

Where questions come from matters more than how many there are. Take them from support tickets, search logs and the questions users actually asked. Questions written by the person who wrote the corpus test the corpus against itself, and they always score better than reality.

Label what actually answers the question, not everything on the topic. Every extra label lowers recall for a retriever that was right, and the metric then rewards a retriever that returns broad, vague context.

## What this evaluation cannot tell you

Four things sit outside what it measures:

**Whether these numbers transfer.** The corpus is ten documents with hand-written 4-dimensional vectors. That makes semantic search unrealistically strong: the question vectors were written to point at the topics the articles cover, with none of the noise a real model produces. Vector search scoring 0.875 here says nothing about your corpus. Run this harness on your data, with your model, before you rely on any of these figures.

**Whether the answer was right.** Recall@k is a ceiling, not a guarantee. A model handed the correct document can still contradict it. Faithfulness and answer quality need their own evaluation, with the generated answer in the loop.

**Whether it is fast enough, or cheap enough.** Neither appears in any of these metrics. The floored hybrid retriever runs two index scans and a fusion for every query.

**Whether the labels are right.** The golden set is a human artefact and it goes stale as the corpus changes. While `array<record<article>>` stops a label pointing at a non-article, nothing stops a label pointing at an article that has since been rewritten to say something else.

## From this set to yours

1. \1.

   Replace the corpus with your documents, and `DIMENSION 4` with your model's dimensionality.
2. \2.

   Write 30 questions from your support log, and label them by hand.
3. \3.

   Keep `fn::retrieve_*`, the three metrics and `fn::evaluate` as they are. They only assume a ranked list of record IDs.
4. \4.

   Put the run in CI, on a fixed `k` that matches how many documents your prompt actually takes.

Then use it on the decision from [lesson 09](https://surrealdb.com/learn/ai/chunking-strategies). Chunk your corpus two ways into two tables, point the retriever at each, and let recall@k choose. That gives you a number for the choice, which is more than any chunking guide can offer.

At a glance

**You are on**

Chapter 10 of 11

**Chapters**

11

**Format**

Text, with queries you can run

**Runs in**

SurrealDB Studio, in the browser

**Cost**

Free

**Certificate**

On completion

[Next: Graph RAG beyond one hop](https://surrealdb.com/learn/ai/graph-rag-multi-hop)

## Continue

### [Chunking strategies](https://surrealdb.com/learn/ai/chunking-strategies)

Previous

### [Graph RAG beyond one hop](https://surrealdb.com/learn/ai/graph-rag-multi-hop)

Next lesson

THE PLATFORM

## Everything an application and its agents know. Five surfaces, one engine.

Database

Document, graph, vector, time-series and relational in one engine.

![Five data models as dotted tiles: documents, graph, vector, time-series and relational](https://surrealdb.com/assets/static/platform-database.DUdumYDz.avif)

Read more

[Database](https://surrealdb.com/surrealdb)

Agent Memory

What an agent learns, with its source and its time, in the same engine.

![A timeline of remembered facts, each with its source](https://surrealdb.com/assets/static/platform-agent-memory.B4RNjvbX.avif)

Read more

[Agent Memory](https://surrealdb.com/agent-memory)

Cloud

Managed clusters in the regions you choose, scaled on demand.

![Clusters in three regions on a world map, each running or scaling](https://surrealdb.com/assets/static/platform-cloud.--QZnaVi.avif)

Read more

[Cloud](https://surrealdb.com/cloud)

Studio

Query, explore and design the schema from the browser.

![A SurrealQL query in Studio and the schema graph under it](https://surrealdb.com/assets/static/platform-studio.7ykBFNLq.avif)

Read more

[Studio](https://surrealdb.com/studio)

MCP

Every model that speaks MCP reaches the database and the memory directly.

![Three models connected through MCP to the database and Agent Memory](https://surrealdb.com/assets/static/platform-mcp.D_oH0_wm.avif)

Read more

[MCP](https://surrealdb.com/mcp)

IN PRODUCTION

## Trusted at scale. Samsung, Nvidia, Verizon, Tencent, and Walmart run on SurrealDB.

14,000+

Developers building on SurrealDB Cloud

4M+

Developers building on SurrealDB worldwide

FROM THE TEAMS

> SurrealDB gives us a foundation where we can unify semantic search, knowledge graphs, and AI-driven decision making without stitching together multiple systems. Collapsing responsibility into SurrealDB has become our default engineering posture.

*Justin Foley*

VP of Engineering, Later

```json
{"@context":"https://schema.org","@type":"Course","name":"SurrealDB for AI Engineers","description":"An eleven-lesson course from calling an LLM API to an agent that retrieves by meaning, remembers what it learned, writes its own queries, and can prove its retrieval works - all against one database.","url":"https://surrealdb.com/learn/ai","inLanguage":"en","isAccessibleForFree":false,"provider":{"@type":"Organization","name":"SurrealDB","url":"https://surrealdb.com"},"hasPart":[{"@type":"LearningResource","name":"SurrealDB for AI Engineers","url":"https://surrealdb.com/learn/ai"},{"@type":"LearningResource","name":"AI foundations","url":"https://surrealdb.com/learn/ai/ai-foundations"},{"@type":"LearningResource","name":"Vector embeddings and search","url":"https://surrealdb.com/learn/ai/vector-embeddings"},{"@type":"LearningResource","name":"Full-text search and BM25","url":"https://surrealdb.com/learn/ai/fulltext-search-bm25"},{"@type":"LearningResource","name":"Building a RAG knowledge base","url":"https://surrealdb.com/learn/ai/rag-knowledge-base"},{"@type":"LearningResource","name":"Hybrid search and reranking","url":"https://surrealdb.com/learn/ai/hybrid-search-reranking"},{"@type":"LearningResource","name":"Building an agent memory store","url":"https://surrealdb.com/learn/ai/agent-memory-store"},{"@type":"LearningResource","name":"Text-to-SurQL: agentic prompt engineering","url":"https://surrealdb.com/learn/ai/text-to-surrealql"},{"@type":"LearningResource","name":"Many sources and agents, one context layer","url":"https://surrealdb.com/learn/ai/multi-source-context-layer"},{"@type":"LearningResource","name":"Chunking strategies","url":"https://surrealdb.com/learn/ai/chunking-strategies"},{"@type":"LearningResource","name":"Evaluating retrieval quality","url":"https://surrealdb.com/learn/ai/evaluating-retrieval"},{"@type":"LearningResource","name":"Graph RAG beyond one hop","url":"https://surrealdb.com/learn/ai/graph-rag-multi-hop"}]}
```

```json
{"@context":"https://schema.org","@type":"LearningResource","name":"Evaluating retrieval quality","description":"Build a golden set as a table, compute recall@k, precision@k and MRR in SurrealQL, and measure three retrievers against each other without leaving the database.","url":"https://surrealdb.com/learn/ai/evaluating-retrieval","learningResourceType":"lesson","isPartOf":{"@type":"Course","name":"SurrealDB for AI Engineers","url":"https://surrealdb.com/learn/ai"},"position":11}
```

```json
{"@context":"https://schema.org","@type":"Organization","@id":"https://surrealdb.com/#organization","name":"SurrealDB","url":"https://surrealdb.com","logo":"https://surrealdb.com/assets/static/logo.BG7_TG2b.svg","description":"SurrealDB is the context and memory layer for AI agents. A multi-model database for documents, graphs, vectors, and time-series.","foundingDate":"2022","legalName":"SurrealDB Ltd","identifier":{"@type":"PropertyValue","propertyID":"GB-COH","value":"13615201"},"address":{"@type":"PostalAddress","streetAddress":"3rd Floor, 1 Ashley Road","addressLocality":"Altrincham","addressRegion":"Cheshire","postalCode":"WA14 2DT","addressCountry":"GB"},"contactPoint":[{"@type":"ContactPoint","contactType":"customer support","email":"support@surrealdb.com","url":"https://surrealdb.com/contact","availableLanguage":"English"},{"@type":"ContactPoint","contactType":"sales","email":"info@surrealdb.com","url":"https://surrealdb.com/contact","availableLanguage":"English"},{"@type":"ContactPoint","contactType":"security","email":"security@surrealdb.com","url":"https://surrealdb.com/.well-known/security.txt","availableLanguage":"English"},{"@type":"ContactPoint","contactType":"legal","email":"legal@surrealdb.com","url":"https://surrealdb.com/legal","availableLanguage":"English"}],"hasCertification":[{"@type":"Certification","name":"SOC 2 Type 2"},{"@type":"Certification","name":"GDPR"},{"@type":"Certification","name":"Cyber Essentials Plus"},{"@type":"Certification","name":"ISO 27001"}],"owns":[{"@type":"SoftwareApplication","name":"SurrealDB","url":"https://surrealdb.com/surrealdb"},{"@type":"SoftwareApplication","name":"Agent Memory","url":"https://surrealdb.com/agent-memory"}],"knowsAbout":["multi-model databases","document databases","graph databases","vector search","time-series databases","SurrealQL","Agent Memory","real-time databases","embedded databases","context layer","graph ontology","distributed database","knowledge graphs","distributed transaction protocols","highly-scalable databases"],"sameAs":["https://www.wikidata.org/wiki/Q124316308","https://github.com/surrealdb/surrealdb","https://twitter.com/surrealdb","https://www.youtube.com/@surrealdb","https://www.linkedin.com/company/surrealdb","https://discord.gg/surrealdb","https://www.reddit.com/r/surrealdb","https://www.instagram.com/surrealdb","https://medium.com/surrealdb","https://dev.to/surrealdb"]}
```

```json
{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://surrealdb.com"},{"@type":"ListItem","position":2,"name":"Learn","item":"https://surrealdb.com/learn"},{"@type":"ListItem","position":3,"name":"Evaluating retrieval","item":"https://surrealdb.com/learn/ai/evaluating-retrieval"}]}
```
