---
title: "Chunking strategies | SurrealDB University"
description: "Chunk the same documents three ways in SurrealDB, measure what each strategy does to retrieval, and expand a small chunk back out to its source at query time."
url: https://surrealdb.com/learn/ai/chunking-strategies
---

![Course content preview](https://surrealdb.com/assets/static/course-ai.LEsq_J_G.avif)

Course chapters

[Back to courses](https://surrealdb.com/learn) [SurrealDB for AI Engineers](https://surrealdb.com/learn/ai) [AI foundations](https://surrealdb.com/learn/ai/ai-foundations) [Vector embeddings and search](https://surrealdb.com/learn/ai/vector-embeddings) [Full-text search and BM25](https://surrealdb.com/learn/ai/fulltext-search-bm25) [Building a RAG knowledge base](https://surrealdb.com/learn/ai/rag-knowledge-base) [Hybrid search and reranking](https://surrealdb.com/learn/ai/hybrid-search-reranking) [Building an agent memory store](https://surrealdb.com/learn/ai/agent-memory-store) [Text-to-SurQL: agentic prompt engineering](https://surrealdb.com/learn/ai/text-to-surrealql) [Many sources and agents, one context layer](https://surrealdb.com/learn/ai/multi-source-context-layer) [Chunking strategies](https://surrealdb.com/learn/ai/chunking-strategies) [Evaluating retrieval quality](https://surrealdb.com/learn/ai/evaluating-retrieval) [Graph RAG beyond one hop](https://surrealdb.com/learn/ai/graph-rag-multi-hop) Certificate Pending

# Chunking strategies

**SurrealQL functions used here for the first time** 

- [`array::clump`](https://surrealdb.com/docs/reference/query-language/functions/database-functions/array#arrayclump) / `.clump()` - cuts an array into fixed-size pieces
- [`array::intersect`](https://surrealdb.com/docs/reference/query-language/functions/database-functions/array#arrayintersect) / `.intersect()` - the values two arrays share, deduplicated
- [`vector::magnitude`](https://surrealdb.com/docs/reference/query-language/functions/database-functions/vector#vectormagnitude) - the length of a vector, and `0` for the zero vector
- [`vector::normalize`](https://surrealdb.com/docs/reference/query-language/functions/database-functions/vector#vectornormalize) - the same vector scaled to length 1
- [`string::words`](https://surrealdb.com/docs/reference/query-language/functions/database-functions/string#stringwords) / `.words()` - splits text on whitespace into an array of words
- [`array::range`](https://surrealdb.com/docs/reference/query-language/functions/database-functions/array#arrayrange) - an array of consecutive integers, used here to number the chunks
- [`type::record`](https://surrealdb.com/docs/reference/query-language/functions/database-functions/type#typerecord) - builds a record id from a table name and an id
- [`record::id`](https://surrealdb.com/docs/reference/query-language/functions/database-functions/record#recordid) / `.id()` - the id half of a record id, without the table name

## Why does chunk size decide retrieval quality?

Because one chunk gets one vector. Everything the chunk says has to be averaged into a single point, so a chunk covering three topics points at none of them, and a chunk cut mid-procedure retrieves the half you did not need. Retrieval quality is capped by the split, and no amount of index tuning lifts that cap.

Both halves of that are visible before any corpus arrives. The first needs a stand-in for an embedding model, and a crude one will do: three topics, each defined by a handful of keywords, and a count of how many of each topic's words a piece of text uses.

```surql
DEFINE ANALYZER OVERWRITE toy
    TOKENIZERS blank, class
    FILTERS lowercase, ascii, snowball(english);

-- A stand-in for an embedding model: counts each topic's keywords in the text
-- and normalises the result.
DEFINE FUNCTION OVERWRITE fn::embed3($text: string) -> array<float> {
    LET $tokens = search::analyze('toy', $text);
    LET $topics = [
        'interlingue occidental auxiliary constructed language grammar vocabulary',
        'xenophon anabasis greek mercenaries persia march historian retreat',
        'venus atmosphere carbon dioxide sulfuric clouds pressure planet'
    ];
    LET $counts = $topics.map(|$topic|
        <float> $tokens.intersect(search::analyze('toy', $topic)).len());
    RETURN IF vector::magnitude($counts) = 0f { [0f, 0f, 0f] } ELSE { vector::normalize($counts) };
};
```

A real model learns its dimensions from text and gives you hundreds of them. This example uses only three for the sake of demonstration, in which each one has a perfect score of 1 for one item in the array and 0 for the rest.

```surql
LET $ling = 'Interlingue, first published as Occidental, is a constructed auxiliary language with a regular grammar and a vocabulary drawn from Romance roots.';
LET $xen = 'Xenophon was a Greek historian whose Anabasis records the retreat of ten thousand mercenaries from Persia.';
LET $ven = 'The atmosphere of Venus is carbon dioxide under crushing pressure, wrapped in sulfuric clouds.';

RETURN { one_topic: fn::embed3($ling), all_three: fn::embed3($ling + ' ' + $xen + ' ' + $ven) };
```

A chunk about one subject is a unit vector on that subject's dimension. A chunk about all three is `1 / sqrt(3)` on each, which is the problem: nothing meaningful sits between a constructed language, a Greek historian and the chemistry of Venus, so the average is not a compromise between them. It is a point that describes none of them.

Output

```surql
{
    one_topic: [1f, 0f, 0f],
    all_three: [0.5773502691896257f, 0.5773502691896257f, 0.5773502691896257f]
}
```

The cost of that lands at query time. `0.577` is also what the mixed chunk scores against a question about Interlingue, and against a question about Xenophon, and against a question about Venus, because a query vector on any one dimension meets the same `0.577` there. A chunk about one subject scores `1` against its own question. The mixed chunk loses to a focused one every time, on every question, including the ones it can answer.

The second failure mode needs no embedder at all, since it is legible. `array::clump` cuts an array into fixed-size pieces, which gives a 40-word split with nothing repeated between the chunks:

```surql
LET $confession = 'I signed the transfer order myself, and I signed the audit that cleared it. For nine years every file you traced to Vienna crossed my desk before it reached yours. The agent you have been hunting since the spring is me.';

RETURN $confession.words().clump(40).map(|$chunk| $chunk.join(' '));
```

Fixed-size chunking counts words and cuts as soon as the limit is reached, regardless of whether the chunk is meaningful or not. This example is a particularly egregious one, in which after forty words of confession the actual answer ends up alone in the second chunk.

Output

```surql
[
    'I signed the transfer order myself, and I signed the audit that cleared it. For nine years every file you traced to Vienna crossed my desk before it reached yours. The agent you have been hunting since the spring is',
    'me.'
]
```

The first chunk is also the one a search for "who was the double agent" would rank first, because it holds Vienna, the nine years and the hunt. A model handed that chunk has everything except the answer, but will pick it as the answer anyway.

Overlap is the usual response to this, and the fixed-size splitter defined later in this lesson takes an overlap argument for that reason. It moves the boundary rather than removing it: whatever the overlap, some sentence somewhere still straddles a cut, and no word count knows which sentence mattered.

The lessons before this one took chunks as given. [Lesson 04](https://surrealdb.com/learn/ai/rag-knowledge-base) seeded eight documents that were already the right size, which is exactly the assumption a real corpus breaks, such as when your input is a 40-page PDF, a wiki with nested headings, or a support thread with six replies.

Five strategies are in common use, and they differ in what they agree to respect:

| Strategy | Cuts on | Costs |
| --- | --- | --- |
| **Fixed-size** | A word or token count, and nothing else | Cuts mid-sentence, as above |
| **Recursive** | Paragraph breaks first, then sentences, then words, splitting further only when a piece is still too big | Needs separators the document actually uses |
| **Semantic** | A drop in similarity between neighbouring sentences, so the cut lands where the subject changes | An embedding call per sentence before you have chunked anything |
| **Document-structure** | Headings, list items, table rows, code fences | Only as good as the document's own markup |
| **Sliding window** | A fixed count again, but with each chunk overlapping the last | Stores the overlapping text more than once |

Our post [What chunking strategies exist and how to choose one](https://surrealdb.com/blog/what-chunk-strategies-exist-and-how-to-choose-one) goes through each in more detail. That post lists the strategies; this lesson measures them, splitting one corpus three ways in the database and showing what each split costs, in numbers you can reproduce.

## Step 1 - start a server

This lesson uses the HNSW index and the `<|K, EF|>` operator from [lesson 02](https://surrealdb.com/learn/ai/vector-embeddings), and the analyzer from [lesson 03](https://surrealdb.com/learn/ai/fulltext-search-bm25). The new part is what happens to the text before either of them sees it.

An in-memory instance in its own terminal is authenticated and throwaway:

```bash
surreal start --user root --pass secret
```

Three files go into it, and each one is linked where it is first used: [`schema.surql`](https://gist.github.com/martinschaer/1cb6fca0634068e5b70f926a5ceed201) for the chunk table, the stand-in embedder and the two splitters, [`seed.surql`](https://gist.github.com/martinschaer/1febc8e96b6771fa36a976cb97ad9b5f) for three documents chunked three ways, and [`queries.surql`](https://gist.github.com/martinschaer/62d3c2299e3f83af9b7166896c163fd7) for the comparison.

## Step 2 - define the schema

### A stand-in for an embedding model

While the other lessons in this course used handwritten toy vectors, that will not work here. The whole question is what the splitter did to the text, so a vector chosen by hand would already contain the answer. The vector has to be derived from whatever the splitter produced.

So the embedder is a SurrealQL function, built the same way as the three-topic one above and with one more topic. Each entry in `$topics` is a handful of keywords standing for one subject the corpus covers. The function counts how many of each subject's keywords the text uses, which gives four numbers, then normalises them so only the proportions survive. A chunk that talks about one subject comes out pointing at that subject; a chunk that talks about two comes out pointing between them.

The keyword lists exist for this lesson only. A real embedding model is handed no vocabulary: it learns its dimensions from the text it was trained on, and those dimensions carry no names a reader would recognise. Counting keywords imitates that crudely, and it is here because the arithmetic stays visible and the same text always gives the same vector, so every number below can be checked by hand.

```surql
DEFINE ANALYZER OVERWRITE toy
    TOKENIZERS blank, class
    FILTERS lowercase, ascii, snowball(english);

-- The same stand-in over this corpus's four topics. The embedding field's
-- VALUE clause calls it on every write.
DEFINE FUNCTION OVERWRITE fn::embed($text: string) -> array<float> {
    LET $tokens = search::analyze('toy', $text);
    LET $topics = [
        'install installation download binary package prerequisites setup',
        'tuning index memory throughput latency benchmark cache',
        'error timeout restart logs failure debug crash',
        'credentials token rotate secret permissions access key'
    ];
    LET $counts = $topics.map(|$topic|
        <float> $tokens.intersect(search::analyze('toy', $topic)).len());
    RETURN IF vector::magnitude($counts) = 0f {
        [0f, 0f, 0f, 0f]
    } ELSE {
        vector::normalize($counts)
    };
};
```

`DEFINE ANALYZER` is the same statement [lesson 03](https://surrealdb.com/learn/ai/fulltext-search-bm25) attached to a full-text index, and `search::analyze` is the function that lesson used to look at the tokens it produces. An analyzer is a named text-processing pipeline and nothing more, so attaching one to an index is something you can do with it rather than a requirement. This one is attached to nothing: the embedder calls it directly through `search::analyze`, on the topic keywords and the chunk text alike. Running both through the same stemmer is why "rotating" in a chunk matches "rotate" in a topic.

A real model reads meaning rather than keywords and returns hundreds of dimensions. What this stand-in shares with it is the property the lesson turns on: one vector per chunk, so the more ground a chunk covers, the less it points anywhere.

One difference matters here: a chunk that hits no keyword embeds to the zero vector, which cosine distance cannot rank. A real model never returns one. The `IF` guard makes that case visible rather than silent.

Another distance measure does not help in this case. Cosine divides by the vector's magnitude, so a zero vector gives `NaN`; Euclidean, Manhattan and Chebyshev are all perfectly defined there and return `1.0` for a unit-length query. That is the worse outcome of the two. A chunk that means nothing then scores as slightly further away than a chunk that is genuinely half relevant, which measured `0.894`, so it competes for a place in the prompt instead of announcing that it cannot be ranked. `NaN` at least refuses to pretend.

### Two splitters

Fixed-size chunking with overlap goes into a function:

```surql
-- Splits text into $size-word chunks, repeating $overlap words between one
-- chunk and the next.
DEFINE FUNCTION OVERWRITE fn::chunk_fixed(
    $text: string,
    $size: int,
    $overlap: int
) -> array<string> {
    LET $words  = $text.words();
    LET $step   = $size - $overlap;
    LET $total  = $words.len();
    LET $extra  = <int> math::ceil(<float> math::max([$total - $size, 0]) / $step);
    LET $chunks = $extra + 1;
    -- Binding a closure to a name before it is used keeps the RETURN line short
    -- enough to read, and the same closure can then be called more than once.
    LET $take   = |$i: int| $words[($i * $step)..math::min([$i * $step + $size, $total])].join(' ');
    RETURN array::range(0, $chunks).map($take).filter(|$chunk| $chunk != '');
};
```

The overlap is what stops a sentence that straddles a boundary from being lost to both chunks. The range indexing has a catch of its own: the third argument of `array::slice` behaves as an end index rather than a length, so `$words[$from..$to]` is the form to trust.

The `<float>` cast is what makes the arithmetic work: integer division truncates before `math::ceil` has anything to round up.

The subtraction is what stops the function over-counting. `math::ceil($total / $step)` asks how many `$step`-sized advances fit in the whole text, but a window is `$size` words wide, not `$step` words wide, so a window near the end can already reach `$total` before that count runs out. With this lesson's own numbers - 194 words, `$size` 60, `$overlap` 12, so `$step` 48 - windows start at 0, 48, 96 and 144, and the one at 144 already covers through word 194, the end of the text. The old formula still asked for a fifth window at 192, and its two words duplicated the tail of the fourth exactly. `$total - $size` asks the question the loop actually needs answered: how many more windows are needed once the first `$size`-word window is placed, so a window that already reaches the end is the last one generated rather than the second-to-last.

Document-structure chunking is shorter, because the document already did the work:

```surql
-- Splits text on its '## ' headings, so the document's own structure decides
-- where the chunks fall.
DEFINE FUNCTION OVERWRITE fn::chunk_sections($text: string) -> array<string> {
    RETURN $text.split('## ')
        .map(|$section| $section.trim())
        .filter(|$section| $section != '');
};
```

### The two tables

```surql
DEFINE TABLE OVERWRITE source SCHEMAFULL;
DEFINE FIELD OVERWRITE title ON source TYPE string;
DEFINE FIELD OVERWRITE body  ON source TYPE string;

DEFINE TABLE OVERWRITE chunk SCHEMAFULL;
DEFINE FIELD OVERWRITE source   ON chunk TYPE record<source>;
DEFINE FIELD OVERWRITE strategy ON chunk TYPE string
    ASSERT $value IN ['whole', 'section', 'fixed'];
DEFINE FIELD OVERWRITE ord      ON chunk TYPE int;
DEFINE FIELD OVERWRITE text     ON chunk TYPE string;

DEFINE FIELD OVERWRITE embedding ON chunk TYPE array<float>
    VALUE fn::embed($this.text)
    ASSERT $value.len() = 4;

DEFINE INDEX OVERWRITE hnsw_chunk ON chunk
    FIELDS embedding
    HNSW DIMENSION 4
    DIST COSINE;

DEFINE INDEX OVERWRITE idx_strategy ON chunk FIELDS strategy;
DEFINE INDEX OVERWRITE idx_position  ON chunk FIELDS source, strategy, ord;
```

That schema makes two choices that the rest of the lesson depends on.

`source` is a **record link**, not a graph edge. A chunk belongs to exactly one source and the connection carries nothing of its own, so a link is the simpler of the two, and `$best.source.body` reads the parent with no join. ([Lesson 11](https://surrealdb.com/learn/ai/graph-rag-multi-hop) uses edges for relationships that do carry their own fields, such as the `kind` on a citation.)

The `VALUE` clause makes `embedding` **derived on write**: whatever the splitter produced, the vector is recomputed from it every time the record is written, so chunking and embedding are one statement and an edited chunk cannot keep a stale vector. In a normal application write, that is all you need.

Load the schema:

```bash
surreal import --endpoint http://localhost:8000 \
  --user root --pass secret --ns ai --db chunking schema.surql
```

## Step 3 - chunk the corpus

The corpus is three documents. The interesting one is the operations handbook for Nimbus, an invented product, whose three sections are about three unrelated things: installing the platform, tuning index builds, and rotating credentials.

Each strategy is a `FOR` loop over the sources. This is the ingest pipeline, and it lives in the database:

```surql
-- Strategy 1 - `whole`: no chunking. The baseline.
FOR $s IN (SELECT id, body FROM source) {
    CREATE type::record('chunk', $s.id.id() + '_whole_0') SET
        source = $s.id, strategy = 'whole', ord = 0, text = $s.body;
};

-- Strategy 2 - `section`: one chunk per heading.
FOR $s IN (SELECT id, body FROM source) {
    LET $sections = fn::chunk_sections($s.body);
    FOR $i IN array::range(0, $sections.len()) {
        CREATE type::record('chunk', $s.id.id() + '_section_' + <string> $i) SET
            source = $s.id, strategy = 'section', ord = $i, text = $sections[$i];
    };
};

-- Strategy 3 - `fixed`: 60 words, 12 of overlap, blind to the headings.
FOR $s IN (SELECT id, body FROM source) {
    LET $chunks = fn::chunk_fixed($s.body, 60, 12);
    FOR $i IN array::range(0, $chunks.len()) {
        CREATE type::record('chunk', $s.id.id() + '_fixed_' + <string> $i) SET
            source = $s.id, strategy = 'fixed', ord = $i, text = $chunks[$i];
    };
};
```

Real pipelines count tokens rather than words, with numbers like 512 and 50. Words keep the output readable here, and the arithmetic is the same.

Record IDs are built from the source, the strategy and the position, so every chunk has a stable, readable ID: `chunk:handbook_section_2`.

```bash
surreal import --endpoint http://localhost:8000 \
  --user root --pass secret --ns ai --db chunking seed.surql
```

## Step 4 - fill in the embeddings

[`seed.surql`](https://gist.github.com/martinschaer/1febc8e96b6771fa36a976cb97ad9b5f) never mentions embeddings. The `CREATE` statements set `text`; the vector is the schema's job, from the `VALUE fn::embed($this.text)` clause on the `embedding` field back in step 2.

And right after the import, there are no vectors:

```bash
surreal sql --endpoint ws://localhost:8000 --user root --pass secret \
  --ns ai --db chunking <<< 'SELECT VALUE embedding FROM chunk LIMIT 1;'
```

Output

```surql
[[NONE]]
```

The `NONE` output is the import doing its job. `surreal import` runs the file in **bulk-load mode**, which turns field processing off along with events and live queries. The `OPTION IMPORT` line at the top of [`seed.surql`](https://gist.github.com/martinschaer/1febc8e96b6771fa36a976cb97ad9b5f) asks for the same mode explicitly, so the file behaves that way whichever route it takes in.

Deferring the vectors is the reason to want that mode. With the `VALUE` clause live, loading 200,000 chunks means 200,000 embedding calls interleaved with the writes, and 200,000 incremental insertions into an HNSW graph that is still growing underneath them. If you land the text first and compute the vectors in one pass afterwards, the index is built once over a settled table.

So this step is part of the load, not a repair to it. One statement, run through `surreal sql` and so outside bulk-load mode, materialises every vector:

```bash
surreal sql --endpoint ws://localhost:8000 --user root --pass secret \
  --ns ai --db chunking <<< 'UPDATE chunk RETURN NONE;'
```

`UPDATE` with no `SET` re-runs the field definitions, which is exactly what a backfill needs. Worth remembering in the other direction too: any bulk import into a table with `VALUE` fields lands without them, so the backfill belongs in the loading script rather than in the incident that follows.

## Step 5 - compare the three strategies

```bash
surreal sql --endpoint ws://localhost:8000 \
  --user root --pass secret --ns ai --db chunking --pretty < queries.surql
```

Every query uses the same question, embedded by the same function that embedded the chunks:

```surql
LET $question = "how do I rotate the service credentials";
LET $q = fn::embed($question);   -- [0f, 0f, 0f, 1f]
```

Both sides have to use the same model. A query vector from a different model is a point in a different space, and every distance computed against it is meaningless.

### What the splitters produced

```surql
SELECT strategy, count() AS chunks,
    math::mean(text.words().len()).round() AS mean_words,
    math::sum(text.words().len()) AS words_stored
FROM chunk GROUP BY strategy;
```

Output

```surql
[
    { strategy: 'fixed',   chunks: 7, mean_words: 51f, words_stored: 359 },
    { strategy: 'section', chunks: 6, mean_words: 51f, words_stored: 305 },
    { strategy: 'whole',   chunks: 3, mean_words: 104f, words_stored: 311 }
]
```

The same three documents and 311 words come out as three different sets of chunks. The `words_stored` field is where overlap shows up as duplication, and 12 words of overlap on a 60-word chunk is why `fixed` stores 359 words of a 311-word corpus. Those 48 duplicated words are stored twice and embedded twice, and a sliding window duplicates far more than that.

### Dilution, in numbers

Here is the handbook as one chunk, next to the same handbook as three:

```surql
SELECT strategy, ord, embedding, text.slice(0, 30) AS starts
FROM chunk
WHERE source = source:handbook AND strategy IN ['whole', 'section']
ORDER BY strategy, ord;
```

Output

```surql
[
    { strategy: 'section', ord: 0, embedding: [1f, 0f, 0f, 0f], starts: 'Installing Nimbus\n\nDownload th' },
    { strategy: 'section', ord: 1, embedding: [0f, 1f, 0f, 0f], starts: 'Tuning index builds\n\nIndex bui' },
    { strategy: 'section', ord: 2, embedding: [0f, 0f, 0f, 1f], starts: 'Rotating service credentials\n\n' },
    { strategy: 'whole',   ord: 0, embedding: [0.577f, 0.577f, 0f, 0.577f], starts: '## Installing Nimbus\n\nDownload' }
]
```

Split by heading, each section is a unit vector on its own axis. Kept whole, the same text is `0.577` on three axes at once: the average of what it says, pointing at nothing in particular. That number is what "dilution" means, and it is why a large chunk loses to a small one on a specific question even when it contains the answer.

### What each strategy can offer

This query runs the same question against each strategy's best chunk, and shows how much text that chunk brings into the prompt:

```surql
SELECT source.title AS source, ord,
    text.words().len() AS words,
    math::round(vector::similarity::cosine(embedding, $q) * 1000) / 1000 AS score
FROM chunk WHERE strategy = 'whole' AND vector::magnitude(embedding) > 0
ORDER BY score DESC LIMIT 1;
```

| strategy | score | words in the prompt |
| --- | --- | --- |
| `whole` | `0.577` | 194 |
| `section` | `1.0` | 66 |
| `fixed` | `1.0` | 50 |

The `vector::magnitude(embedding) > 0` clause is a defensive habit worth keeping regardless: a chunk that hits no topic keyword at all embeds to the zero vector this stand-in can produce, for the reason shown in the previous step, and cosine similarity against it is `NaN`, which sorts ahead of `1.0` in `ORDER BY score DESC` rather than last. These use `vector::similarity::cosine` directly rather than the HNSW index, because a comparison wants every chunk scored rather than the top K. Exact search on a corpus this size is instant; see [lesson 02](https://surrealdb.com/learn/ai/vector-embeddings) for what to do when the corpus reaches a much larger size.

The whole-document strategy retrieves the right document and hands the model 194 words, two thirds of which are about installing and tuning. It also scores lowest on the one question it can answer.

### Where the fixed-size split went wrong

The `fixed` strategy scored a perfect `1.0`, but it is still the wrong thing to send:

```surql
SELECT ord, embedding, text.slice(0, 64) AS starts
FROM chunk
WHERE source = source:handbook AND strategy = 'fixed'
ORDER BY ord;
```

Output

```surql
[
    { ord: 0, embedding: [1f, 0f, 0f, 0f],           starts: '## Installing Nimbus\n\nDownload the release binary for your ' },
    { ord: 1, embedding: [0.141f, 0.99f, 0f, 0f],    starts: 'then start the service. A fresh installation binds to localhost ' },
    { ord: 2, embedding: [0f, 0.8f, 0f, 0.6f],       starts: 'improves without costing latency on the read path. Benchmark eve' },
    { ord: 3, embedding: [0f, 0f, 0f, 1f],           starts: 'deploy it alongside the old one, and only then revoke the previo' },
]
```

Two failures are visible in that output.

Chunk 2 straddles a heading. Its vector is `[0, 0.8, 0, 0.6]`, and those positions are the four topic axes `fn::embed` counts against: install, tuning, errors, credentials. So the chunk reads as `0.8` tuning and `0.6` credentials at once, because the 60-word window ran out in the middle of the tuning section and kept going into the credentials one. It is the chunk a question about either topic half-matches.

Chunk 3 is the one that scored `1.0`, and it starts with "deploy it alongside the old one". The first two steps of the rotation procedure, "issue the new token first", are in chunk 2. Fixed-size chunking retrieved the right topic and cut the instruction in half.

## Step 6 - retrieve small, generate big

Chunk size does not have to be the trade-off. Retrieve the small, precise chunk, and widen the text afterwards, before it reaches the model. Start with the retrieval:

```surql
LET $best = (SELECT id, source, ord, text, vector::distance::knn() AS dist
    FROM chunk
    WHERE strategy = 'fixed' AND embedding <|5, 40|> $q
    ORDER BY dist)[0];
RETURN { chunk: $best.id, dist: $best.dist };
```

Output

```surql
{ chunk: chunk:handbook_fixed_3, dist: 0f }
```

The `strategy = 'fixed'` filter is pushed **into** the index scan rather than applied to its output, so `<|5, 40|>` returns five chunks that already match the predicate:

```surql
EXPLAIN SELECT id FROM chunk WHERE strategy = 'whole' AND embedding <|1, 40|> $q;
```

Output

```surql
"SelectProject [ctx: Db] [projections: id]
    Filter [ctx: Db] [predicate: strategy = 'whole']
        KnnScan [ctx: Db] [index: hnsw_chunk, k: 1, ef: 40, dimension: 4, predicate: strategy = 'whole']"
```

That is what makes typed filters and vector search compose in one query instead of one narrowing the other by accident.

Now comes the expansion. The neighbouring chunk holds the start of the procedure, and `ord` is all you need to find it:

```surql
SELECT ord, text.slice(0, 48) AS starts
FROM chunk
WHERE source = $best.source AND strategy = 'fixed'
    AND ord IN [$best.ord - 1, $best.ord, $best.ord + 1]
ORDER BY ord;
```

Output

```surql
[
    { ord: 2, starts: 'improves without costing latency on the read pat' },
    { ord: 3, starts: 'deploy it alongside the old one, and only then r' }
]
```

Join those two together, though, and the overlap comes back twice: "revoke the previous" appears at the end of chunk 2 and again at the start of chunk 3. Overlap helps at retrieval time, but the duplicate text then has to be removed when neighbouring chunks are joined.

With an overlapping strategy, expand to the parent instead of stitching neighbours:

```surql
RETURN { cite: $best.source.title, passage: $best.source.body };
```

That is the pattern to take away. **Chunk small enough to be found, store a link to something big enough to be useful, and expand at query time.** The retrieval unit and the generation unit need not be the same object. The record link is what lets you separate them.

### Adding a chunk later

Because the `VALUE` clause derives the embedding, a new chunk is one write:

```surql
CREATE chunk:handbook_section_3 SET
    source = source:handbook, strategy = 'section', ord = 3,
    text = "Revoking a leaked key. Revoke the leaked key immediately, issue a replacement token, and audit the access logs for requests that used the old secret.";
SELECT ord, embedding FROM chunk:handbook_section_3;
```

Output

```surql
[
    { ord: 3, embedding: [0f, 0f, 0.243f, 0.97f] }
]
```

The text and its vector are written by one statement, so they cannot disagree.

## Choosing a strategy

The [chunking strategies post](https://surrealdb.com/blog/what-chunk-strategies-exist-and-how-to-choose-one) has the full decision path. The short version, in the terms this lesson measured:

| If your corpus | Start with | Because |
| --- | --- | --- |
| Has headings, or is markdown, HTML or code | document-structure | The author already marked where topics change |
| Is unstructured prose | recursive, then fixed-size with overlap | Sentence boundaries are the cheapest approximation of a topic boundary |
| Mixes unrelated topics with no structure | semantic | Worth the embedding cost when nothing else marks the boundary |
| Needs answers spanning a boundary | any of the above, plus expansion | Step 6, rather than a larger chunk |

Two numbers to hold on to from step 5: the whole-document vector was `0.577` where the section vector was `1.0`, and the fixed-size split cut a three-step procedure across two chunks. Both are the split talking, not the index.

## What to measure

Every measurement in this lesson came from a single question, which makes it a sample of one. A chunk size that wins on "how do I rotate the service credentials" can lose on the next question, and you will not know unless you measure it.

That is [lesson 10](https://surrealdb.com/learn/ai/evaluating-retrieval): a golden set of questions with the answers labelled, recall@k, and the harness to run one retriever against another. Chunk size is the first thing to put through it.

## From the stand-in to a real model

Three changes:

1. \1.

   Replace the body of `fn::embed` with a call to your model, or embed in the application and pass the vector in. [Lesson 02](https://surrealdb.com/learn/ai/vector-embeddings) covers choosing a model.
2. \2.

   Change `DIMENSION 4` and the `$value.len() = 4` assertion to your model's dimensionality.
3. \3.

   Count tokens rather than words in `fn::chunk_fixed`, using your model's tokeniser.

Everything else you can pass on: the splitters, the `source` link, the derived `embedding`, the `ord` window, the parent expansion.

One upgrade is worth trying before anything else. Prepend the source title, and the heading path, to the chunk text before embedding it: `$source.title + ' > ' + $heading + '\n\n' + $chunk`. A chunk that starts mid-sentence gets its context back, and it costs one string concatenation.

At a glance

**You are on**

Chapter 9 of 11

**Chapters**

11

**Format**

Text, with queries you can run

**Runs in**

SurrealDB Studio, in the browser

**Cost**

Free

**Certificate**

On completion

[Next: Evaluating retrieval quality](https://surrealdb.com/learn/ai/evaluating-retrieval)

## Continue

### [Many sources and agents, one context layer](https://surrealdb.com/learn/ai/multi-source-context-layer)

Previous

### [Evaluating retrieval quality](https://surrealdb.com/learn/ai/evaluating-retrieval)

Next lesson

THE PLATFORM

## Everything an application and its agents know. Five surfaces, one engine.

Database

Document, graph, vector, time-series and relational in one engine.

![Five data models as dotted tiles: documents, graph, vector, time-series and relational](https://surrealdb.com/assets/static/platform-database.DUdumYDz.avif)

Read more

[Database](https://surrealdb.com/surrealdb)

Agent Memory

What an agent learns, with its source and its time, in the same engine.

![A timeline of remembered facts, each with its source](https://surrealdb.com/assets/static/platform-agent-memory.B4RNjvbX.avif)

Read more

[Agent Memory](https://surrealdb.com/agent-memory)

Cloud

Managed clusters in the regions you choose, scaled on demand.

![Clusters in three regions on a world map, each running or scaling](https://surrealdb.com/assets/static/platform-cloud.--QZnaVi.avif)

Read more

[Cloud](https://surrealdb.com/cloud)

Studio

Query, explore and design the schema from the browser.

![A SurrealQL query in Studio and the schema graph under it](https://surrealdb.com/assets/static/platform-studio.7ykBFNLq.avif)

Read more

[Studio](https://surrealdb.com/studio)

MCP

Every model that speaks MCP reaches the database and the memory directly.

![Three models connected through MCP to the database and Agent Memory](https://surrealdb.com/assets/static/platform-mcp.D_oH0_wm.avif)

Read more

[MCP](https://surrealdb.com/mcp)

IN PRODUCTION

## Trusted at scale. Samsung, Nvidia, Verizon, Tencent, and Walmart run on SurrealDB.

14,000+

Developers building on SurrealDB Cloud

4M+

Developers building on SurrealDB worldwide

FROM THE TEAMS

> SurrealDB gives us a foundation where we can unify semantic search, knowledge graphs, and AI-driven decision making without stitching together multiple systems. Collapsing responsibility into SurrealDB has become our default engineering posture.

*Justin Foley*

VP of Engineering, Later

```json
{"@context":"https://schema.org","@type":"Course","name":"SurrealDB for AI Engineers","description":"An eleven-lesson course from calling an LLM API to an agent that retrieves by meaning, remembers what it learned, writes its own queries, and can prove its retrieval works - all against one database.","url":"https://surrealdb.com/learn/ai","inLanguage":"en","isAccessibleForFree":false,"provider":{"@type":"Organization","name":"SurrealDB","url":"https://surrealdb.com"},"hasPart":[{"@type":"LearningResource","name":"SurrealDB for AI Engineers","url":"https://surrealdb.com/learn/ai"},{"@type":"LearningResource","name":"AI foundations","url":"https://surrealdb.com/learn/ai/ai-foundations"},{"@type":"LearningResource","name":"Vector embeddings and search","url":"https://surrealdb.com/learn/ai/vector-embeddings"},{"@type":"LearningResource","name":"Full-text search and BM25","url":"https://surrealdb.com/learn/ai/fulltext-search-bm25"},{"@type":"LearningResource","name":"Building a RAG knowledge base","url":"https://surrealdb.com/learn/ai/rag-knowledge-base"},{"@type":"LearningResource","name":"Hybrid search and reranking","url":"https://surrealdb.com/learn/ai/hybrid-search-reranking"},{"@type":"LearningResource","name":"Building an agent memory store","url":"https://surrealdb.com/learn/ai/agent-memory-store"},{"@type":"LearningResource","name":"Text-to-SurQL: agentic prompt engineering","url":"https://surrealdb.com/learn/ai/text-to-surrealql"},{"@type":"LearningResource","name":"Many sources and agents, one context layer","url":"https://surrealdb.com/learn/ai/multi-source-context-layer"},{"@type":"LearningResource","name":"Chunking strategies","url":"https://surrealdb.com/learn/ai/chunking-strategies"},{"@type":"LearningResource","name":"Evaluating retrieval quality","url":"https://surrealdb.com/learn/ai/evaluating-retrieval"},{"@type":"LearningResource","name":"Graph RAG beyond one hop","url":"https://surrealdb.com/learn/ai/graph-rag-multi-hop"}]}
```

```json
{"@context":"https://schema.org","@type":"LearningResource","name":"Chunking strategies","description":"Chunk the same documents three ways in SurrealDB, measure what each strategy does to retrieval, and expand a small chunk back out to its source at query time.","url":"https://surrealdb.com/learn/ai/chunking-strategies","learningResourceType":"lesson","isPartOf":{"@type":"Course","name":"SurrealDB for AI Engineers","url":"https://surrealdb.com/learn/ai"},"position":10}
```

```json
{"@context":"https://schema.org","@type":"Organization","@id":"https://surrealdb.com/#organization","name":"SurrealDB","url":"https://surrealdb.com","logo":"https://surrealdb.com/assets/static/logo.BG7_TG2b.svg","description":"SurrealDB is the context and memory layer for AI agents. A multi-model database for documents, graphs, vectors, and time-series.","foundingDate":"2022","legalName":"SurrealDB Ltd","identifier":{"@type":"PropertyValue","propertyID":"GB-COH","value":"13615201"},"address":{"@type":"PostalAddress","streetAddress":"3rd Floor, 1 Ashley Road","addressLocality":"Altrincham","addressRegion":"Cheshire","postalCode":"WA14 2DT","addressCountry":"GB"},"contactPoint":[{"@type":"ContactPoint","contactType":"customer support","email":"support@surrealdb.com","url":"https://surrealdb.com/contact","availableLanguage":"English"},{"@type":"ContactPoint","contactType":"sales","email":"info@surrealdb.com","url":"https://surrealdb.com/contact","availableLanguage":"English"},{"@type":"ContactPoint","contactType":"security","email":"security@surrealdb.com","url":"https://surrealdb.com/.well-known/security.txt","availableLanguage":"English"},{"@type":"ContactPoint","contactType":"legal","email":"legal@surrealdb.com","url":"https://surrealdb.com/legal","availableLanguage":"English"}],"hasCertification":[{"@type":"Certification","name":"SOC 2 Type 2"},{"@type":"Certification","name":"GDPR"},{"@type":"Certification","name":"Cyber Essentials Plus"},{"@type":"Certification","name":"ISO 27001"}],"owns":[{"@type":"SoftwareApplication","name":"SurrealDB","url":"https://surrealdb.com/surrealdb"},{"@type":"SoftwareApplication","name":"Agent Memory","url":"https://surrealdb.com/agent-memory"}],"knowsAbout":["multi-model databases","document databases","graph databases","vector search","time-series databases","SurrealQL","Agent Memory","real-time databases","embedded databases","context layer","graph ontology","distributed database","knowledge graphs","distributed transaction protocols","highly-scalable databases"],"sameAs":["https://www.wikidata.org/wiki/Q124316308","https://github.com/surrealdb/surrealdb","https://twitter.com/surrealdb","https://www.youtube.com/@surrealdb","https://www.linkedin.com/company/surrealdb","https://discord.gg/surrealdb","https://www.reddit.com/r/surrealdb","https://www.instagram.com/surrealdb","https://medium.com/surrealdb","https://dev.to/surrealdb"]}
```

```json
{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://surrealdb.com"},{"@type":"ListItem","position":2,"name":"Learn","item":"https://surrealdb.com/learn"},{"@type":"ListItem","position":3,"name":"Chunking strategies","item":"https://surrealdb.com/learn/ai/chunking-strategies"}]}
```
