Skip to content
Course content preview

Chunking strategies

SurrealQL functions used here for the first time

  • array::clump / .clump() - cuts an array into fixed-size pieces

  • array::intersect / .intersect() - the values two arrays share, deduplicated

  • vector::magnitude - the length of a vector, and 0 for the zero vector

  • vector::normalize - the same vector scaled to length 1

  • string::words / .words() - splits text on whitespace into an array of words

  • array::range - an array of consecutive integers, used here to number the chunks

  • type::record - builds a record id from a table name and an id

  • record::id / .id() - the id half of a record id, without the table name

Because one chunk gets one vector. Everything the chunk says has to be averaged into a single point, so a chunk covering three topics points at none of them, and a chunk cut mid-procedure retrieves the half you did not need. Retrieval quality is capped by the split, and no amount of index tuning lifts that cap.

Both halves of that are visible before any corpus arrives. The first needs a stand-in for an embedding model, and a crude one will do: three topics, each defined by a handful of keywords, and a count of how many of each topic's words a piece of text uses.

DEFINE ANALYZER OVERWRITE toy
    TOKENIZERS blank, class
    FILTERS lowercase, ascii, snowball(english);

-- A stand-in for an embedding model: counts each topic's keywords in the text
-- and normalises the result.
DEFINE FUNCTION OVERWRITE fn::embed3($text: string) -> array<float> {
    LET $tokens = search::analyze('toy', $text);
    LET $topics = [
        'interlingue occidental auxiliary constructed language grammar vocabulary',
        'xenophon anabasis greek mercenaries persia march historian retreat',
        'venus atmosphere carbon dioxide sulfuric clouds pressure planet'
    ];
    LET $counts = $topics.map(|$topic|
        <float> $tokens.intersect(search::analyze('toy', $topic)).len());
    RETURN IF vector::magnitude($counts) = 0f { [0f, 0f, 0f] } ELSE { vector::normalize($counts) };
};

A real model learns its dimensions from text and gives you hundreds of them. This example uses only three for the sake of demonstration, in which each one has a perfect score of 1 for one item in the array and 0 for the rest.

LET $ling = 'Interlingue, first published as Occidental, is a constructed auxiliary language with a regular grammar and a vocabulary drawn from Romance roots.';
LET $xen = 'Xenophon was a Greek historian whose Anabasis records the retreat of ten thousand mercenaries from Persia.';
LET $ven = 'The atmosphere of Venus is carbon dioxide under crushing pressure, wrapped in sulfuric clouds.';

RETURN { one_topic: fn::embed3($ling), all_three: fn::embed3($ling + ' ' + $xen + ' ' + $ven) };

A chunk about one subject is a unit vector on that subject's dimension. A chunk about all three is 1 / sqrt(3) on each, which is the problem: nothing meaningful sits between a constructed language, a Greek historian and the chemistry of Venus, so the average is not a compromise between them. It is a point that describes none of them.

Output
{
    one_topic: [1f, 0f, 0f],
    all_three: [0.5773502691896257f, 0.5773502691896257f, 0.5773502691896257f]
}

The cost of that lands at query time. 0.577 is also what the mixed chunk scores against a question about Interlingue, and against a question about Xenophon, and against a question about Venus, because a query vector on any one dimension meets the same 0.577 there. A chunk about one subject scores 1 against its own question. The mixed chunk loses to a focused one every time, on every question, including the ones it can answer.

The second failure mode needs no embedder at all, since it is legible. array::clump cuts an array into fixed-size pieces, which gives a 40-word split with nothing repeated between the chunks:

LET $confession = 'I signed the transfer order myself, and I signed the audit that cleared it. For nine years every file you traced to Vienna crossed my desk before it reached yours. The agent you have been hunting since the spring is me.';

RETURN $confession.words().clump(40).map(|$chunk| $chunk.join(' '));

Fixed-size chunking counts words and cuts as soon as the limit is reached, regardless of whether the chunk is meaningful or not. This example is a particularly egregious one, in which after forty words of confession the actual answer ends up alone in the second chunk.

Output
[
    'I signed the transfer order myself, and I signed the audit that cleared it. For nine years every file you traced to Vienna crossed my desk before it reached yours. The agent you have been hunting since the spring is',
    'me.'
]

The first chunk is also the one a search for "who was the double agent" would rank first, because it holds Vienna, the nine years and the hunt. A model handed that chunk has everything except the answer, but will pick it as the answer anyway.

Overlap is the usual response to this, and the fixed-size splitter defined later in this lesson takes an overlap argument for that reason. It moves the boundary rather than removing it: whatever the overlap, some sentence somewhere still straddles a cut, and no word count knows which sentence mattered.

The lessons before this one took chunks as given. Lesson 04 seeded eight documents that were already the right size, which is exactly the assumption a real corpus breaks, such as when your input is a 40-page PDF, a wiki with nested headings, or a support thread with six replies.

Five strategies are in common use, and they differ in what they agree to respect:

StrategyCuts onCosts
Fixed-sizeA word or token count, and nothing elseCuts mid-sentence, as above
RecursiveParagraph breaks first, then sentences, then words, splitting further only when a piece is still too bigNeeds separators the document actually uses
SemanticA drop in similarity between neighbouring sentences, so the cut lands where the subject changesAn embedding call per sentence before you have chunked anything
Document-structureHeadings, list items, table rows, code fencesOnly as good as the document's own markup
Sliding windowA fixed count again, but with each chunk overlapping the lastStores the overlapping text more than once

Our post What chunking strategies exist and how to choose one goes through each in more detail. That post lists the strategies; this lesson measures them, splitting one corpus three ways in the database and showing what each split costs, in numbers you can reproduce.

This lesson uses the HNSW index and the <|K, EF|> operator from lesson 02, and the analyzer from lesson 03. The new part is what happens to the text before either of them sees it.

An in-memory instance in its own terminal is authenticated and throwaway:

surreal start --user root --pass secret

Three files go into it, and each one is linked where it is first used: schema.surql for the chunk table, the stand-in embedder and the two splitters, seed.surql for three documents chunked three ways, and queries.surql for the comparison.

While the other lessons in this course used handwritten toy vectors, that will not work here. The whole question is what the splitter did to the text, so a vector chosen by hand would already contain the answer. The vector has to be derived from whatever the splitter produced.

So the embedder is a SurrealQL function, built the same way as the three-topic one above and with one more topic. Each entry in $topics is a handful of keywords standing for one subject the corpus covers. The function counts how many of each subject's keywords the text uses, which gives four numbers, then normalises them so only the proportions survive. A chunk that talks about one subject comes out pointing at that subject; a chunk that talks about two comes out pointing between them.

The keyword lists exist for this lesson only. A real embedding model is handed no vocabulary: it learns its dimensions from the text it was trained on, and those dimensions carry no names a reader would recognise. Counting keywords imitates that crudely, and it is here because the arithmetic stays visible and the same text always gives the same vector, so every number below can be checked by hand.

DEFINE ANALYZER OVERWRITE toy
    TOKENIZERS blank, class
    FILTERS lowercase, ascii, snowball(english);

-- The same stand-in over this corpus's four topics. The embedding field's
-- VALUE clause calls it on every write.
DEFINE FUNCTION OVERWRITE fn::embed($text: string) -> array<float> {
    LET $tokens = search::analyze('toy', $text);
    LET $topics = [
        'install installation download binary package prerequisites setup',
        'tuning index memory throughput latency benchmark cache',
        'error timeout restart logs failure debug crash',
        'credentials token rotate secret permissions access key'
    ];
    LET $counts = $topics.map(|$topic|
        <float> $tokens.intersect(search::analyze('toy', $topic)).len());
    RETURN IF vector::magnitude($counts) = 0f {
        [0f, 0f, 0f, 0f]
    } ELSE {
        vector::normalize($counts)
    };
};

DEFINE ANALYZER is the same statement lesson 03 attached to a full-text index, and search::analyze is the function that lesson used to look at the tokens it produces. An analyzer is a named text-processing pipeline and nothing more, so attaching one to an index is something you can do with it rather than a requirement. This one is attached to nothing: the embedder calls it directly through search::analyze, on the topic keywords and the chunk text alike. Running both through the same stemmer is why "rotating" in a chunk matches "rotate" in a topic.

A real model reads meaning rather than keywords and returns hundreds of dimensions. What this stand-in shares with it is the property the lesson turns on: one vector per chunk, so the more ground a chunk covers, the less it points anywhere.

One difference matters here: a chunk that hits no keyword embeds to the zero vector, which cosine distance cannot rank. A real model never returns one. The IF guard makes that case visible rather than silent.

Another distance measure does not help in this case. Cosine divides by the vector's magnitude, so a zero vector gives NaN; Euclidean, Manhattan and Chebyshev are all perfectly defined there and return 1.0 for a unit-length query. That is the worse outcome of the two. A chunk that means nothing then scores as slightly further away than a chunk that is genuinely half relevant, which measured 0.894, so it competes for a place in the prompt instead of announcing that it cannot be ranked. NaN at least refuses to pretend.

Fixed-size chunking with overlap goes into a function:

-- Splits text into $size-word chunks, repeating $overlap words between one
-- chunk and the next.
DEFINE FUNCTION OVERWRITE fn::chunk_fixed(
    $text: string,
    $size: int,
    $overlap: int
) -> array<string> {
    LET $words  = $text.words();
    LET $step   = $size - $overlap;
    LET $total  = $words.len();
    LET $extra  = <int> math::ceil(<float> math::max([$total - $size, 0]) / $step);
    LET $chunks = $extra + 1;
    -- Binding a closure to a name before it is used keeps the RETURN line short
    -- enough to read, and the same closure can then be called more than once.
    LET $take   = |$i: int| $words[($i * $step)..math::min([$i * $step + $size, $total])].join(' ');
    RETURN array::range(0, $chunks).map($take).filter(|$chunk| $chunk != '');
};

The overlap is what stops a sentence that straddles a boundary from being lost to both chunks. The range indexing has a catch of its own: the third argument of array::slice behaves as an end index rather than a length, so $words[$from..$to] is the form to trust.

The <float> cast is what makes the arithmetic work: integer division truncates before math::ceil has anything to round up.

The subtraction is what stops the function over-counting. math::ceil($total / $step) asks how many $step-sized advances fit in the whole text, but a window is $size words wide, not $step words wide, so a window near the end can already reach $total before that count runs out. With this lesson's own numbers - 194 words, $size 60, $overlap 12, so $step 48 - windows start at 0, 48, 96 and 144, and the one at 144 already covers through word 194, the end of the text. The old formula still asked for a fifth window at 192, and its two words duplicated the tail of the fourth exactly. $total - $size asks the question the loop actually needs answered: how many more windows are needed once the first $size-word window is placed, so a window that already reaches the end is the last one generated rather than the second-to-last.

Document-structure chunking is shorter, because the document already did the work:

-- Splits text on its '## ' headings, so the document's own structure decides
-- where the chunks fall.
DEFINE FUNCTION OVERWRITE fn::chunk_sections($text: string) -> array<string> {
    RETURN $text.split('## ')
        .map(|$section| $section.trim())
        .filter(|$section| $section != '');
};
DEFINE TABLE OVERWRITE source SCHEMAFULL;
DEFINE FIELD OVERWRITE title ON source TYPE string;
DEFINE FIELD OVERWRITE body  ON source TYPE string;

DEFINE TABLE OVERWRITE chunk SCHEMAFULL;
DEFINE FIELD OVERWRITE source   ON chunk TYPE record<source>;
DEFINE FIELD OVERWRITE strategy ON chunk TYPE string
    ASSERT $value IN ['whole', 'section', 'fixed'];
DEFINE FIELD OVERWRITE ord      ON chunk TYPE int;
DEFINE FIELD OVERWRITE text     ON chunk TYPE string;

DEFINE FIELD OVERWRITE embedding ON chunk TYPE array<float>
    VALUE fn::embed($this.text)
    ASSERT $value.len() = 4;

DEFINE INDEX OVERWRITE hnsw_chunk ON chunk
    FIELDS embedding
    HNSW DIMENSION 4
    DIST COSINE;

DEFINE INDEX OVERWRITE idx_strategy ON chunk FIELDS strategy;
DEFINE INDEX OVERWRITE idx_position  ON chunk FIELDS source, strategy, ord;

That schema makes two choices that the rest of the lesson depends on.

source is a record link, not a graph edge. A chunk belongs to exactly one source and the connection carries nothing of its own, so a link is the simpler of the two, and $best.source.body reads the parent with no join. (Lesson 11 uses edges for relationships that do carry their own fields, such as the kind on a citation.)

The VALUE clause makes embedding derived on write: whatever the splitter produced, the vector is recomputed from it every time the record is written, so chunking and embedding are one statement and an edited chunk cannot keep a stale vector. In a normal application write, that is all you need.

Load the schema:

surreal import --endpoint http://localhost:8000 \
  --user root --pass secret --ns ai --db chunking schema.surql

The corpus is three documents. The interesting one is the operations handbook for Nimbus, an invented product, whose three sections are about three unrelated things: installing the platform, tuning index builds, and rotating credentials.

Each strategy is a FOR loop over the sources. This is the ingest pipeline, and it lives in the database:

-- Strategy 1 - `whole`: no chunking. The baseline.
FOR $s IN (SELECT id, body FROM source) {
    CREATE type::record('chunk', $s.id.id() + '_whole_0') SET
        source = $s.id, strategy = 'whole', ord = 0, text = $s.body;
};

-- Strategy 2 - `section`: one chunk per heading.
FOR $s IN (SELECT id, body FROM source) {
    LET $sections = fn::chunk_sections($s.body);
    FOR $i IN array::range(0, $sections.len()) {
        CREATE type::record('chunk', $s.id.id() + '_section_' + <string> $i) SET
            source = $s.id, strategy = 'section', ord = $i, text = $sections[$i];
    };
};

-- Strategy 3 - `fixed`: 60 words, 12 of overlap, blind to the headings.
FOR $s IN (SELECT id, body FROM source) {
    LET $chunks = fn::chunk_fixed($s.body, 60, 12);
    FOR $i IN array::range(0, $chunks.len()) {
        CREATE type::record('chunk', $s.id.id() + '_fixed_' + <string> $i) SET
            source = $s.id, strategy = 'fixed', ord = $i, text = $chunks[$i];
    };
};

Real pipelines count tokens rather than words, with numbers like 512 and 50. Words keep the output readable here, and the arithmetic is the same.

Record IDs are built from the source, the strategy and the position, so every chunk has a stable, readable ID: chunk:handbook_section_2.

surreal import --endpoint http://localhost:8000 \
  --user root --pass secret --ns ai --db chunking seed.surql

seed.surql never mentions embeddings. The CREATE statements set text; the vector is the schema's job, from the VALUE fn::embed($this.text) clause on the embedding field back in step 2.

And right after the import, there are no vectors:

surreal sql --endpoint ws://localhost:8000 --user root --pass secret \
  --ns ai --db chunking <<< 'SELECT VALUE embedding FROM chunk LIMIT 1;'
Output
[[NONE]]

The NONE output is the import doing its job. surreal import runs the file in bulk-load mode, which turns field processing off along with events and live queries. The OPTION IMPORT line at the top of seed.surql asks for the same mode explicitly, so the file behaves that way whichever route it takes in.

Deferring the vectors is the reason to want that mode. With the VALUE clause live, loading 200,000 chunks means 200,000 embedding calls interleaved with the writes, and 200,000 incremental insertions into an HNSW graph that is still growing underneath them. If you land the text first and compute the vectors in one pass afterwards, the index is built once over a settled table.

So this step is part of the load, not a repair to it. One statement, run through surreal sql and so outside bulk-load mode, materialises every vector:

surreal sql --endpoint ws://localhost:8000 --user root --pass secret \
  --ns ai --db chunking <<< 'UPDATE chunk RETURN NONE;'

UPDATE with no SET re-runs the field definitions, which is exactly what a backfill needs. Worth remembering in the other direction too: any bulk import into a table with VALUE fields lands without them, so the backfill belongs in the loading script rather than in the incident that follows.

surreal sql --endpoint ws://localhost:8000 \
  --user root --pass secret --ns ai --db chunking --pretty < queries.surql

Every query uses the same question, embedded by the same function that embedded the chunks:

LET $question = "how do I rotate the service credentials";
LET $q = fn::embed($question);   -- [0f, 0f, 0f, 1f]

Both sides have to use the same model. A query vector from a different model is a point in a different space, and every distance computed against it is meaningless.

SELECT strategy, count() AS chunks,
    math::mean(text.words().len()).round() AS mean_words,
    math::sum(text.words().len()) AS words_stored
FROM chunk GROUP BY strategy;
Output
[
    { strategy: 'fixed',   chunks: 7, mean_words: 51f, words_stored: 359 },
    { strategy: 'section', chunks: 6, mean_words: 51f, words_stored: 305 },
    { strategy: 'whole',   chunks: 3, mean_words: 104f, words_stored: 311 }
]

The same three documents and 311 words come out as three different sets of chunks. The words_stored field is where overlap shows up as duplication, and 12 words of overlap on a 60-word chunk is why fixed stores 359 words of a 311-word corpus. Those 48 duplicated words are stored twice and embedded twice, and a sliding window duplicates far more than that.

Here is the handbook as one chunk, next to the same handbook as three:

SELECT strategy, ord, embedding, text.slice(0, 30) AS starts
FROM chunk
WHERE source = source:handbook AND strategy IN ['whole', 'section']
ORDER BY strategy, ord;
Output
[
    { strategy: 'section', ord: 0, embedding: [1f, 0f, 0f, 0f], starts: 'Installing Nimbus\n\nDownload th' },
    { strategy: 'section', ord: 1, embedding: [0f, 1f, 0f, 0f], starts: 'Tuning index builds\n\nIndex bui' },
    { strategy: 'section', ord: 2, embedding: [0f, 0f, 0f, 1f], starts: 'Rotating service credentials\n\n' },
    { strategy: 'whole',   ord: 0, embedding: [0.577f, 0.577f, 0f, 0.577f], starts: '## Installing Nimbus\n\nDownload' }
]

Split by heading, each section is a unit vector on its own axis. Kept whole, the same text is 0.577 on three axes at once: the average of what it says, pointing at nothing in particular. That number is what "dilution" means, and it is why a large chunk loses to a small one on a specific question even when it contains the answer.

This query runs the same question against each strategy's best chunk, and shows how much text that chunk brings into the prompt:

SELECT source.title AS source, ord,
    text.words().len() AS words,
    math::round(vector::similarity::cosine(embedding, $q) * 1000) / 1000 AS score
FROM chunk WHERE strategy = 'whole' AND vector::magnitude(embedding) > 0
ORDER BY score DESC LIMIT 1;
strategyscorewords in the prompt
whole0.577194
section1.066
fixed1.050

The vector::magnitude(embedding) > 0 clause is a defensive habit worth keeping regardless: a chunk that hits no topic keyword at all embeds to the zero vector this stand-in can produce, for the reason shown in the previous step, and cosine similarity against it is NaN, which sorts ahead of 1.0 in ORDER BY score DESC rather than last. These use vector::similarity::cosine directly rather than the HNSW index, because a comparison wants every chunk scored rather than the top K. Exact search on a corpus this size is instant; see lesson 02 for what to do when the corpus reaches a much larger size.

The whole-document strategy retrieves the right document and hands the model 194 words, two thirds of which are about installing and tuning. It also scores lowest on the one question it can answer.

The fixed strategy scored a perfect 1.0, but it is still the wrong thing to send:

SELECT ord, embedding, text.slice(0, 64) AS starts
FROM chunk
WHERE source = source:handbook AND strategy = 'fixed'
ORDER BY ord;
Output
[
    { ord: 0, embedding: [1f, 0f, 0f, 0f],           starts: '## Installing Nimbus\n\nDownload the release binary for your ' },
    { ord: 1, embedding: [0.141f, 0.99f, 0f, 0f],    starts: 'then start the service. A fresh installation binds to localhost ' },
    { ord: 2, embedding: [0f, 0.8f, 0f, 0.6f],       starts: 'improves without costing latency on the read path. Benchmark eve' },
    { ord: 3, embedding: [0f, 0f, 0f, 1f],           starts: 'deploy it alongside the old one, and only then revoke the previo' },
]

Two failures are visible in that output.

Chunk 2 straddles a heading. Its vector is [0, 0.8, 0, 0.6], and those positions are the four topic axes fn::embed counts against: install, tuning, errors, credentials. So the chunk reads as 0.8 tuning and 0.6 credentials at once, because the 60-word window ran out in the middle of the tuning section and kept going into the credentials one. It is the chunk a question about either topic half-matches.

Chunk 3 is the one that scored 1.0, and it starts with "deploy it alongside the old one". The first two steps of the rotation procedure, "issue the new token first", are in chunk 2. Fixed-size chunking retrieved the right topic and cut the instruction in half.

Chunk size does not have to be the trade-off. Retrieve the small, precise chunk, and widen the text afterwards, before it reaches the model. Start with the retrieval:

LET $best = (SELECT id, source, ord, text, vector::distance::knn() AS dist
    FROM chunk
    WHERE strategy = 'fixed' AND embedding <|5, 40|> $q
    ORDER BY dist)[0];
RETURN { chunk: $best.id, dist: $best.dist };
Output
{ chunk: chunk:handbook_fixed_3, dist: 0f }

The strategy = 'fixed' filter is pushed into the index scan rather than applied to its output, so <|5, 40|> returns five chunks that already match the predicate:

EXPLAIN SELECT id FROM chunk WHERE strategy = 'whole' AND embedding <|1, 40|> $q;
Output
"SelectProject [ctx: Db] [projections: id]
    Filter [ctx: Db] [predicate: strategy = 'whole']
        KnnScan [ctx: Db] [index: hnsw_chunk, k: 1, ef: 40, dimension: 4, predicate: strategy = 'whole']"

That is what makes typed filters and vector search compose in one query instead of one narrowing the other by accident.

Now comes the expansion. The neighbouring chunk holds the start of the procedure, and ord is all you need to find it:

SELECT ord, text.slice(0, 48) AS starts
FROM chunk
WHERE source = $best.source AND strategy = 'fixed'
    AND ord IN [$best.ord - 1, $best.ord, $best.ord + 1]
ORDER BY ord;
Output
[
    { ord: 2, starts: 'improves without costing latency on the read pat' },
    { ord: 3, starts: 'deploy it alongside the old one, and only then r' }
]

Join those two together, though, and the overlap comes back twice: "revoke the previous" appears at the end of chunk 2 and again at the start of chunk 3. Overlap helps at retrieval time, but the duplicate text then has to be removed when neighbouring chunks are joined.

With an overlapping strategy, expand to the parent instead of stitching neighbours:

RETURN { cite: $best.source.title, passage: $best.source.body };

That is the pattern to take away. Chunk small enough to be found, store a link to something big enough to be useful, and expand at query time. The retrieval unit and the generation unit need not be the same object. The record link is what lets you separate them.

Because the VALUE clause derives the embedding, a new chunk is one write:

CREATE chunk:handbook_section_3 SET
    source = source:handbook, strategy = 'section', ord = 3,
    text = "Revoking a leaked key. Revoke the leaked key immediately, issue a replacement token, and audit the access logs for requests that used the old secret.";
SELECT ord, embedding FROM chunk:handbook_section_3;
Output
[
    { ord: 3, embedding: [0f, 0f, 0.243f, 0.97f] }
]

The text and its vector are written by one statement, so they cannot disagree.

The chunking strategies post has the full decision path. The short version, in the terms this lesson measured:

If your corpusStart withBecause
Has headings, or is markdown, HTML or codedocument-structureThe author already marked where topics change
Is unstructured proserecursive, then fixed-size with overlapSentence boundaries are the cheapest approximation of a topic boundary
Mixes unrelated topics with no structuresemanticWorth the embedding cost when nothing else marks the boundary
Needs answers spanning a boundaryany of the above, plus expansionStep 6, rather than a larger chunk

Two numbers to hold on to from step 5: the whole-document vector was 0.577 where the section vector was 1.0, and the fixed-size split cut a three-step procedure across two chunks. Both are the split talking, not the index.

Every measurement in this lesson came from a single question, which makes it a sample of one. A chunk size that wins on "how do I rotate the service credentials" can lose on the next question, and you will not know unless you measure it.

That is lesson 10: a golden set of questions with the answers labelled, recall@k, and the harness to run one retriever against another. Chunk size is the first thing to put through it.

Three changes:

  1. 1.

    Replace the body of fn::embed with a call to your model, or embed in the application and pass the vector in. Lesson 02 covers choosing a model.

  2. 2.

    Change DIMENSION 4 and the $value.len() = 4 assertion to your model's dimensionality.

  3. 3.

    Count tokens rather than words in fn::chunk_fixed, using your model's tokeniser.

Everything else you can pass on: the splitters, the source link, the derived embedding, the ord window, the parent expansion.

One upgrade is worth trying before anything else. Prepend the source title, and the heading path, to the chunk text before embedding it: $source.title + ' > ' + $heading + '\n\n' + $chunk. A chunk that starts mid-sentence gets its context back, and it costs one string concatenation.

THE PLATFORM

Everything an application and its agents know. Five surfaces, one engine.

IN PRODUCTION

Trusted at scale. Samsung, Nvidia, Verizon, Tencent, and Walmart run on SurrealDB.

14,000+

Developers building on SurrealDB Cloud

4M+

Developers building on SurrealDB worldwide

FROM THE TEAMS

SurrealDB gives us a foundation where we can unify semantic search, knowledge graphs, and AI-driven decision making without stitching together multiple systems. Collapsing responsibility into SurrealDB has become our default engineering posture.
Justin Foley

VP of Engineering, Later

SurrealDB

The context and memory layer for AI agents

Database. Graphs, vectors, documents and relational data in one engine, in a single ACID transaction.
Agent Memory. Connects and retrieves context wherever your data lives, every fact carrying its source.
Cloud. Fully managed, in the cloud provider and region you choose.

Explore with AI

Copyright © 2026 SurrealDB Ltd. Registered in England and Wales. Company no. 13615201

Registered address: 3rd Floor 1 Ashley Road, Altrincham, Cheshire, WA14 2DT, United Kingdom

Trading address: Huckletree Oxford Circus, 213 Oxford Street, London, W1D 2LG, United Kingdom