
Chunking strategies
SurrealQL functions used here for the first time
array::clump/.clump()- cuts an array into fixed-size piecesarray::intersect/.intersect()- the values two arrays share, deduplicatedvector::magnitude- the length of a vector, and0for the zero vectorvector::normalize- the same vector scaled to length 1string::words/.words()- splits text on whitespace into an array of wordsarray::range- an array of consecutive integers, used here to number the chunkstype::record- builds a record id from a table name and an idrecord::id/.id()- the id half of a record id, without the table name
Why does chunk size decide retrieval quality?
Because one chunk gets one vector. Everything the chunk says has to be averaged into a single point, so a chunk covering three topics points at none of them, and a chunk cut mid-procedure retrieves the half you did not need. Retrieval quality is capped by the split, and no amount of index tuning lifts that cap.
Both halves of that are visible before any corpus arrives. The first needs a stand-in for an embedding model, and a crude one will do: three topics, each defined by a handful of keywords, and a count of how many of each topic's words a piece of text uses.
DEFINE ANALYZER OVERWRITE toy
TOKENIZERS blank, class
FILTERS lowercase, ascii, snowball(english);
-- A stand-in for an embedding model: counts each topic's keywords in the text
-- and normalises the result.
DEFINE FUNCTION OVERWRITE fn::embed3($text: string) -> array<float> {
LET $tokens = search::analyze('toy', $text);
LET $topics = [
'interlingue occidental auxiliary constructed language grammar vocabulary',
'xenophon anabasis greek mercenaries persia march historian retreat',
'venus atmosphere carbon dioxide sulfuric clouds pressure planet'
];
LET $counts = $topics.map(|$topic|
<float> $tokens.intersect(search::analyze('toy', $topic)).len());
RETURN IF vector::magnitude($counts) = 0f { [0f, 0f, 0f] } ELSE { vector::normalize($counts) };
};A real model learns its dimensions from text and gives you hundreds of them. This example uses only three for the sake of demonstration, in which each one has a perfect score of 1 for one item in the array and 0 for the rest.
LET $ling = 'Interlingue, first published as Occidental, is a constructed auxiliary language with a regular grammar and a vocabulary drawn from Romance roots.';
LET $xen = 'Xenophon was a Greek historian whose Anabasis records the retreat of ten thousand mercenaries from Persia.';
LET $ven = 'The atmosphere of Venus is carbon dioxide under crushing pressure, wrapped in sulfuric clouds.';
RETURN { one_topic: fn::embed3($ling), all_three: fn::embed3($ling + ' ' + $xen + ' ' + $ven) };A chunk about one subject is a unit vector on that subject's dimension. A chunk about all three is 1 / sqrt(3) on each, which is the problem: nothing meaningful sits between a constructed language, a Greek historian and the chemistry of Venus, so the average is not a compromise between them. It is a point that describes none of them.
{
one_topic: [1f, 0f, 0f],
all_three: [0.5773502691896257f, 0.5773502691896257f, 0.5773502691896257f]
}The cost of that lands at query time. 0.577 is also what the mixed chunk scores against a question about Interlingue, and against a question about Xenophon, and against a question about Venus, because a query vector on any one dimension meets the same 0.577 there. A chunk about one subject scores 1 against its own question. The mixed chunk loses to a focused one every time, on every question, including the ones it can answer.
The second failure mode needs no embedder at all, since it is legible. array::clump cuts an array into fixed-size pieces, which gives a 40-word split with nothing repeated between the chunks:
LET $confession = 'I signed the transfer order myself, and I signed the audit that cleared it. For nine years every file you traced to Vienna crossed my desk before it reached yours. The agent you have been hunting since the spring is me.';
RETURN $confession.words().clump(40).map(|$chunk| $chunk.join(' '));Fixed-size chunking counts words and cuts as soon as the limit is reached, regardless of whether the chunk is meaningful or not. This example is a particularly egregious one, in which after forty words of confession the actual answer ends up alone in the second chunk.
[
'I signed the transfer order myself, and I signed the audit that cleared it. For nine years every file you traced to Vienna crossed my desk before it reached yours. The agent you have been hunting since the spring is',
'me.'
]The first chunk is also the one a search for "who was the double agent" would rank first, because it holds Vienna, the nine years and the hunt. A model handed that chunk has everything except the answer, but will pick it as the answer anyway.
Overlap is the usual response to this, and the fixed-size splitter defined later in this lesson takes an overlap argument for that reason. It moves the boundary rather than removing it: whatever the overlap, some sentence somewhere still straddles a cut, and no word count knows which sentence mattered.
The lessons before this one took chunks as given. Lesson 04 seeded eight documents that were already the right size, which is exactly the assumption a real corpus breaks, such as when your input is a 40-page PDF, a wiki with nested headings, or a support thread with six replies.
Five strategies are in common use, and they differ in what they agree to respect:
| Strategy | Cuts on | Costs |
|---|---|---|
| Fixed-size | A word or token count, and nothing else | Cuts mid-sentence, as above |
| Recursive | Paragraph breaks first, then sentences, then words, splitting further only when a piece is still too big | Needs separators the document actually uses |
| Semantic | A drop in similarity between neighbouring sentences, so the cut lands where the subject changes | An embedding call per sentence before you have chunked anything |
| Document-structure | Headings, list items, table rows, code fences | Only as good as the document's own markup |
| Sliding window | A fixed count again, but with each chunk overlapping the last | Stores the overlapping text more than once |
Our post What chunking strategies exist and how to choose one goes through each in more detail. That post lists the strategies; this lesson measures them, splitting one corpus three ways in the database and showing what each split costs, in numbers you can reproduce.
Step 1 - start a server
This lesson uses the HNSW index and the <|K, EF|> operator from lesson 02, and the analyzer from lesson 03. The new part is what happens to the text before either of them sees it.
An in-memory instance in its own terminal is authenticated and throwaway:
surreal start --user root --pass secretThree files go into it, and each one is linked where it is first used: schema.surql for the chunk table, the stand-in embedder and the two splitters, seed.surql for three documents chunked three ways, and queries.surql for the comparison.
Step 2 - define the schema
A stand-in for an embedding model
While the other lessons in this course used handwritten toy vectors, that will not work here. The whole question is what the splitter did to the text, so a vector chosen by hand would already contain the answer. The vector has to be derived from whatever the splitter produced.
So the embedder is a SurrealQL function, built the same way as the three-topic one above and with one more topic. Each entry in $topics is a handful of keywords standing for one subject the corpus covers. The function counts how many of each subject's keywords the text uses, which gives four numbers, then normalises them so only the proportions survive. A chunk that talks about one subject comes out pointing at that subject; a chunk that talks about two comes out pointing between them.
The keyword lists exist for this lesson only. A real embedding model is handed no vocabulary: it learns its dimensions from the text it was trained on, and those dimensions carry no names a reader would recognise. Counting keywords imitates that crudely, and it is here because the arithmetic stays visible and the same text always gives the same vector, so every number below can be checked by hand.
DEFINE ANALYZER OVERWRITE toy
TOKENIZERS blank, class
FILTERS lowercase, ascii, snowball(english);
-- The same stand-in over this corpus's four topics. The embedding field's
-- VALUE clause calls it on every write.
DEFINE FUNCTION OVERWRITE fn::embed($text: string) -> array<float> {
LET $tokens = search::analyze('toy', $text);
LET $topics = [
'install installation download binary package prerequisites setup',
'tuning index memory throughput latency benchmark cache',
'error timeout restart logs failure debug crash',
'credentials token rotate secret permissions access key'
];
LET $counts = $topics.map(|$topic|
<float> $tokens.intersect(search::analyze('toy', $topic)).len());
RETURN IF vector::magnitude($counts) = 0f {
[0f, 0f, 0f, 0f]
} ELSE {
vector::normalize($counts)
};
};DEFINE ANALYZER is the same statement lesson 03 attached to a full-text index, and search::analyze is the function that lesson used to look at the tokens it produces. An analyzer is a named text-processing pipeline and nothing more, so attaching one to an index is something you can do with it rather than a requirement. This one is attached to nothing: the embedder calls it directly through search::analyze, on the topic keywords and the chunk text alike. Running both through the same stemmer is why "rotating" in a chunk matches "rotate" in a topic.
A real model reads meaning rather than keywords and returns hundreds of dimensions. What this stand-in shares with it is the property the lesson turns on: one vector per chunk, so the more ground a chunk covers, the less it points anywhere.
One difference matters here: a chunk that hits no keyword embeds to the zero vector, which cosine distance cannot rank. A real model never returns one. The IF guard makes that case visible rather than silent.
Another distance measure does not help in this case. Cosine divides by the vector's magnitude, so a zero vector gives NaN; Euclidean, Manhattan and Chebyshev are all perfectly defined there and return 1.0 for a unit-length query. That is the worse outcome of the two. A chunk that means nothing then scores as slightly further away than a chunk that is genuinely half relevant, which measured 0.894, so it competes for a place in the prompt instead of announcing that it cannot be ranked. NaN at least refuses to pretend.
Two splitters
Fixed-size chunking with overlap goes into a function:
-- Splits text into $size-word chunks, repeating $overlap words between one
-- chunk and the next.
DEFINE FUNCTION OVERWRITE fn::chunk_fixed(
$text: string,
$size: int,
$overlap: int
) -> array<string> {
LET $words = $text.words();
LET $step = $size - $overlap;
LET $total = $words.len();
LET $extra = <int> math::ceil(<float> math::max([$total - $size, 0]) / $step);
LET $chunks = $extra + 1;
-- Binding a closure to a name before it is used keeps the RETURN line short
-- enough to read, and the same closure can then be called more than once.
LET $take = |$i: int| $words[($i * $step)..math::min([$i * $step + $size, $total])].join(' ');
RETURN array::range(0, $chunks).map($take).filter(|$chunk| $chunk != '');
};The overlap is what stops a sentence that straddles a boundary from being lost to both chunks. The range indexing has a catch of its own: the third argument of array::slice behaves as an end index rather than a length, so $words[$from..$to] is the form to trust.
The <float> cast is what makes the arithmetic work: integer division truncates before math::ceil has anything to round up.
The subtraction is what stops the function over-counting. math::ceil($total / $step) asks how many $step-sized advances fit in the whole text, but a window is $size words wide, not $step words wide, so a window near the end can already reach $total before that count runs out. With this lesson's own numbers - 194 words, $size 60, $overlap 12, so $step 48 - windows start at 0, 48, 96 and 144, and the one at 144 already covers through word 194, the end of the text. The old formula still asked for a fifth window at 192, and its two words duplicated the tail of the fourth exactly. $total - $size asks the question the loop actually needs answered: how many more windows are needed once the first $size-word window is placed, so a window that already reaches the end is the last one generated rather than the second-to-last.
Document-structure chunking is shorter, because the document already did the work:
-- Splits text on its '## ' headings, so the document's own structure decides
-- where the chunks fall.
DEFINE FUNCTION OVERWRITE fn::chunk_sections($text: string) -> array<string> {
RETURN $text.split('## ')
.map(|$section| $section.trim())
.filter(|$section| $section != '');
};The two tables
DEFINE TABLE OVERWRITE source SCHEMAFULL;
DEFINE FIELD OVERWRITE title ON source TYPE string;
DEFINE FIELD OVERWRITE body ON source TYPE string;
DEFINE TABLE OVERWRITE chunk SCHEMAFULL;
DEFINE FIELD OVERWRITE source ON chunk TYPE record<source>;
DEFINE FIELD OVERWRITE strategy ON chunk TYPE string
ASSERT $value IN ['whole', 'section', 'fixed'];
DEFINE FIELD OVERWRITE ord ON chunk TYPE int;
DEFINE FIELD OVERWRITE text ON chunk TYPE string;
DEFINE FIELD OVERWRITE embedding ON chunk TYPE array<float>
VALUE fn::embed($this.text)
ASSERT $value.len() = 4;
DEFINE INDEX OVERWRITE hnsw_chunk ON chunk
FIELDS embedding
HNSW DIMENSION 4
DIST COSINE;
DEFINE INDEX OVERWRITE idx_strategy ON chunk FIELDS strategy;
DEFINE INDEX OVERWRITE idx_position ON chunk FIELDS source, strategy, ord;That schema makes two choices that the rest of the lesson depends on.
source is a record link, not a graph edge. A chunk belongs to exactly one source and the connection carries nothing of its own, so a link is the simpler of the two, and $best.source.body reads the parent with no join. (Lesson 11 uses edges for relationships that do carry their own fields, such as the kind on a citation.)
The VALUE clause makes embedding derived on write: whatever the splitter produced, the vector is recomputed from it every time the record is written, so chunking and embedding are one statement and an edited chunk cannot keep a stale vector. In a normal application write, that is all you need.
Load the schema:
surreal import --endpoint http://localhost:8000 \
--user root --pass secret --ns ai --db chunking schema.surqlStep 3 - chunk the corpus
The corpus is three documents. The interesting one is the operations handbook for Nimbus, an invented product, whose three sections are about three unrelated things: installing the platform, tuning index builds, and rotating credentials.
Each strategy is a FOR loop over the sources. This is the ingest pipeline, and it lives in the database:
-- Strategy 1 - `whole`: no chunking. The baseline.
FOR $s IN (SELECT id, body FROM source) {
CREATE type::record('chunk', $s.id.id() + '_whole_0') SET
source = $s.id, strategy = 'whole', ord = 0, text = $s.body;
};
-- Strategy 2 - `section`: one chunk per heading.
FOR $s IN (SELECT id, body FROM source) {
LET $sections = fn::chunk_sections($s.body);
FOR $i IN array::range(0, $sections.len()) {
CREATE type::record('chunk', $s.id.id() + '_section_' + <string> $i) SET
source = $s.id, strategy = 'section', ord = $i, text = $sections[$i];
};
};
-- Strategy 3 - `fixed`: 60 words, 12 of overlap, blind to the headings.
FOR $s IN (SELECT id, body FROM source) {
LET $chunks = fn::chunk_fixed($s.body, 60, 12);
FOR $i IN array::range(0, $chunks.len()) {
CREATE type::record('chunk', $s.id.id() + '_fixed_' + <string> $i) SET
source = $s.id, strategy = 'fixed', ord = $i, text = $chunks[$i];
};
};Real pipelines count tokens rather than words, with numbers like 512 and 50. Words keep the output readable here, and the arithmetic is the same.
Record IDs are built from the source, the strategy and the position, so every chunk has a stable, readable ID: chunk:handbook_section_2.
surreal import --endpoint http://localhost:8000 \
--user root --pass secret --ns ai --db chunking seed.surqlStep 4 - fill in the embeddings
seed.surql never mentions embeddings. The CREATE statements set text; the vector is the schema's job, from the VALUE fn::embed($this.text) clause on the embedding field back in step 2.
And right after the import, there are no vectors:
surreal sql --endpoint ws://localhost:8000 --user root --pass secret \
--ns ai --db chunking <<< 'SELECT VALUE embedding FROM chunk LIMIT 1;'[[NONE]]The NONE output is the import doing its job. surreal import runs the file in bulk-load mode, which turns field processing off along with events and live queries. The OPTION IMPORT line at the top of seed.surql asks for the same mode explicitly, so the file behaves that way whichever route it takes in.
Deferring the vectors is the reason to want that mode. With the VALUE clause live, loading 200,000 chunks means 200,000 embedding calls interleaved with the writes, and 200,000 incremental insertions into an HNSW graph that is still growing underneath them. If you land the text first and compute the vectors in one pass afterwards, the index is built once over a settled table.
So this step is part of the load, not a repair to it. One statement, run through surreal sql and so outside bulk-load mode, materialises every vector:
surreal sql --endpoint ws://localhost:8000 --user root --pass secret \
--ns ai --db chunking <<< 'UPDATE chunk RETURN NONE;'UPDATE with no SET re-runs the field definitions, which is exactly what a backfill needs. Worth remembering in the other direction too: any bulk import into a table with VALUE fields lands without them, so the backfill belongs in the loading script rather than in the incident that follows.
Step 5 - compare the three strategies
surreal sql --endpoint ws://localhost:8000 \
--user root --pass secret --ns ai --db chunking --pretty < queries.surqlEvery query uses the same question, embedded by the same function that embedded the chunks:
LET $question = "how do I rotate the service credentials";
LET $q = fn::embed($question); -- [0f, 0f, 0f, 1f]Both sides have to use the same model. A query vector from a different model is a point in a different space, and every distance computed against it is meaningless.
What the splitters produced
SELECT strategy, count() AS chunks,
math::mean(text.words().len()).round() AS mean_words,
math::sum(text.words().len()) AS words_stored
FROM chunk GROUP BY strategy;[
{ strategy: 'fixed', chunks: 7, mean_words: 51f, words_stored: 359 },
{ strategy: 'section', chunks: 6, mean_words: 51f, words_stored: 305 },
{ strategy: 'whole', chunks: 3, mean_words: 104f, words_stored: 311 }
]The same three documents and 311 words come out as three different sets of chunks. The words_stored field is where overlap shows up as duplication, and 12 words of overlap on a 60-word chunk is why fixed stores 359 words of a 311-word corpus. Those 48 duplicated words are stored twice and embedded twice, and a sliding window duplicates far more than that.
Dilution, in numbers
Here is the handbook as one chunk, next to the same handbook as three:
SELECT strategy, ord, embedding, text.slice(0, 30) AS starts
FROM chunk
WHERE source = source:handbook AND strategy IN ['whole', 'section']
ORDER BY strategy, ord;[
{ strategy: 'section', ord: 0, embedding: [1f, 0f, 0f, 0f], starts: 'Installing Nimbus\n\nDownload th' },
{ strategy: 'section', ord: 1, embedding: [0f, 1f, 0f, 0f], starts: 'Tuning index builds\n\nIndex bui' },
{ strategy: 'section', ord: 2, embedding: [0f, 0f, 0f, 1f], starts: 'Rotating service credentials\n\n' },
{ strategy: 'whole', ord: 0, embedding: [0.577f, 0.577f, 0f, 0.577f], starts: '## Installing Nimbus\n\nDownload' }
]Split by heading, each section is a unit vector on its own axis. Kept whole, the same text is 0.577 on three axes at once: the average of what it says, pointing at nothing in particular. That number is what "dilution" means, and it is why a large chunk loses to a small one on a specific question even when it contains the answer.
What each strategy can offer
This query runs the same question against each strategy's best chunk, and shows how much text that chunk brings into the prompt:
SELECT source.title AS source, ord,
text.words().len() AS words,
math::round(vector::similarity::cosine(embedding, $q) * 1000) / 1000 AS score
FROM chunk WHERE strategy = 'whole' AND vector::magnitude(embedding) > 0
ORDER BY score DESC LIMIT 1;| strategy | score | words in the prompt |
|---|---|---|
whole | 0.577 | 194 |
section | 1.0 | 66 |
fixed | 1.0 | 50 |
The vector::magnitude(embedding) > 0 clause is a defensive habit worth keeping regardless: a chunk that hits no topic keyword at all embeds to the zero vector this stand-in can produce, for the reason shown in the previous step, and cosine similarity against it is NaN, which sorts ahead of 1.0 in ORDER BY score DESC rather than last. These use vector::similarity::cosine directly rather than the HNSW index, because a comparison wants every chunk scored rather than the top K. Exact search on a corpus this size is instant; see lesson 02 for what to do when the corpus reaches a much larger size.
The whole-document strategy retrieves the right document and hands the model 194 words, two thirds of which are about installing and tuning. It also scores lowest on the one question it can answer.
Where the fixed-size split went wrong
The fixed strategy scored a perfect 1.0, but it is still the wrong thing to send:
SELECT ord, embedding, text.slice(0, 64) AS starts
FROM chunk
WHERE source = source:handbook AND strategy = 'fixed'
ORDER BY ord;[
{ ord: 0, embedding: [1f, 0f, 0f, 0f], starts: '## Installing Nimbus\n\nDownload the release binary for your ' },
{ ord: 1, embedding: [0.141f, 0.99f, 0f, 0f], starts: 'then start the service. A fresh installation binds to localhost ' },
{ ord: 2, embedding: [0f, 0.8f, 0f, 0.6f], starts: 'improves without costing latency on the read path. Benchmark eve' },
{ ord: 3, embedding: [0f, 0f, 0f, 1f], starts: 'deploy it alongside the old one, and only then revoke the previo' },
]Two failures are visible in that output.
Chunk 2 straddles a heading. Its vector is [0, 0.8, 0, 0.6], and those positions are the four topic axes fn::embed counts against: install, tuning, errors, credentials. So the chunk reads as 0.8 tuning and 0.6 credentials at once, because the 60-word window ran out in the middle of the tuning section and kept going into the credentials one. It is the chunk a question about either topic half-matches.
Chunk 3 is the one that scored 1.0, and it starts with "deploy it alongside the old one". The first two steps of the rotation procedure, "issue the new token first", are in chunk 2. Fixed-size chunking retrieved the right topic and cut the instruction in half.
Step 6 - retrieve small, generate big
Chunk size does not have to be the trade-off. Retrieve the small, precise chunk, and widen the text afterwards, before it reaches the model. Start with the retrieval:
LET $best = (SELECT id, source, ord, text, vector::distance::knn() AS dist
FROM chunk
WHERE strategy = 'fixed' AND embedding <|5, 40|> $q
ORDER BY dist)[0];
RETURN { chunk: $best.id, dist: $best.dist };{ chunk: chunk:handbook_fixed_3, dist: 0f }The strategy = 'fixed' filter is pushed into the index scan rather than applied to its output, so <|5, 40|> returns five chunks that already match the predicate:
EXPLAIN SELECT id FROM chunk WHERE strategy = 'whole' AND embedding <|1, 40|> $q;"SelectProject [ctx: Db] [projections: id]
Filter [ctx: Db] [predicate: strategy = 'whole']
KnnScan [ctx: Db] [index: hnsw_chunk, k: 1, ef: 40, dimension: 4, predicate: strategy = 'whole']"That is what makes typed filters and vector search compose in one query instead of one narrowing the other by accident.
Now comes the expansion. The neighbouring chunk holds the start of the procedure, and ord is all you need to find it:
SELECT ord, text.slice(0, 48) AS starts
FROM chunk
WHERE source = $best.source AND strategy = 'fixed'
AND ord IN [$best.ord - 1, $best.ord, $best.ord + 1]
ORDER BY ord;[
{ ord: 2, starts: 'improves without costing latency on the read pat' },
{ ord: 3, starts: 'deploy it alongside the old one, and only then r' }
]Join those two together, though, and the overlap comes back twice: "revoke the previous" appears at the end of chunk 2 and again at the start of chunk 3. Overlap helps at retrieval time, but the duplicate text then has to be removed when neighbouring chunks are joined.
With an overlapping strategy, expand to the parent instead of stitching neighbours:
RETURN { cite: $best.source.title, passage: $best.source.body };That is the pattern to take away. Chunk small enough to be found, store a link to something big enough to be useful, and expand at query time. The retrieval unit and the generation unit need not be the same object. The record link is what lets you separate them.
Adding a chunk later
Because the VALUE clause derives the embedding, a new chunk is one write:
CREATE chunk:handbook_section_3 SET
source = source:handbook, strategy = 'section', ord = 3,
text = "Revoking a leaked key. Revoke the leaked key immediately, issue a replacement token, and audit the access logs for requests that used the old secret.";
SELECT ord, embedding FROM chunk:handbook_section_3;[
{ ord: 3, embedding: [0f, 0f, 0.243f, 0.97f] }
]The text and its vector are written by one statement, so they cannot disagree.
Choosing a strategy
The chunking strategies post has the full decision path. The short version, in the terms this lesson measured:
| If your corpus | Start with | Because |
|---|---|---|
| Has headings, or is markdown, HTML or code | document-structure | The author already marked where topics change |
| Is unstructured prose | recursive, then fixed-size with overlap | Sentence boundaries are the cheapest approximation of a topic boundary |
| Mixes unrelated topics with no structure | semantic | Worth the embedding cost when nothing else marks the boundary |
| Needs answers spanning a boundary | any of the above, plus expansion | Step 6, rather than a larger chunk |
Two numbers to hold on to from step 5: the whole-document vector was 0.577 where the section vector was 1.0, and the fixed-size split cut a three-step procedure across two chunks. Both are the split talking, not the index.
What to measure
Every measurement in this lesson came from a single question, which makes it a sample of one. A chunk size that wins on "how do I rotate the service credentials" can lose on the next question, and you will not know unless you measure it.
That is lesson 10: a golden set of questions with the answers labelled, recall@k, and the harness to run one retriever against another. Chunk size is the first thing to put through it.
From the stand-in to a real model
Three changes:
- 1.
Replace the body of
fn::embedwith a call to your model, or embed in the application and pass the vector in. Lesson 02 covers choosing a model. - 2.
Change
DIMENSION 4and the$value.len() = 4assertion to your model's dimensionality. - 3.
Count tokens rather than words in
fn::chunk_fixed, using your model's tokeniser.
Everything else you can pass on: the splitters, the source link, the derived embedding, the ord window, the parent expansion.
One upgrade is worth trying before anything else. Prepend the source title, and the heading path, to the chunk text before embedding it: $source.title + ' > ' + $heading + '\n\n' + $chunk. A chunk that starts mid-sentence gets its context back, and it costs one string concatenation.
THE PLATFORM
Everything an application and its agents know. Five surfaces, one engine.
Database
Document, graph, vector, time-series and relational in one engine.

Agent Memory
What an agent learns, with its source and its time, in the same engine.

Cloud
Managed clusters in the regions you choose, scaled on demand.

Studio
Query, explore and design the schema from the browser.

MCP
Every model that speaks MCP reaches the database and the memory directly.

IN PRODUCTION
Trusted at scale. Samsung, Nvidia, Verizon, Tencent, and Walmart run on SurrealDB.
14,000+
Developers building on SurrealDB Cloud
4M+
Developers building on SurrealDB worldwide
FROM THE TEAMS
SurrealDB gives us a foundation where we can unify semantic search, knowledge graphs, and AI-driven decision making without stitching together multiple systems. Collapsing responsibility into SurrealDB has become our default engineering posture.
VP of Engineering, Later