
Graph RAG beyond one hop
SurrealQL functions used here for the first time
array::group/.group()- collapses one level of nesting and removes duplicates, which is what a multi-hop traversal needsarray::complement/.complement()- the values in one array that are not in another
Why is one hop not enough?
Because the document that explains a problem is often not the document that mentions it. If you ask why a vector index is slow to build, the nearest document by meaning is about index parameters. The cause is write amplification during compaction, three citations away, in a document that shares almost no vocabulary with the question.
Similarity ranks documents by how much they look like the question. A graph ranks them by how they connect to the answer. When the cause is a hop away, only the second one finds it, and lesson 04's single ->cites->document reaches exactly one step of it.
Step 1 - a graph to traverse
This lesson assumes lesson 04: the HNSW index, the <|K, EF|> operator, and RELATE. The corpus is lesson 04's, grown to fourteen documents and given a citation chain deep enough to get lost in. Three files go in: schema.surql for the documents, typed citation edges, topic nodes and the HNSW index, seed.surql for fourteen documents, thirteen citations and four topics, and queries.surql for the traversals.
The document table is lesson 04's. What changes is the edge:
DEFINE TABLE OVERWRITE cites SCHEMAFULL TYPE RELATION FROM document TO document;
DEFINE FIELD OVERWRITE kind ON cites TYPE 'background' | 'method' | 'benchmark';
DEFINE FIELD OVERWRITE year ON cites TYPE int;
DEFINE INDEX OVERWRITE idx_cites_kind ON cites FIELDS kind;Lesson 04 used a bare RELATE a->cites->b. A bare edge is fine for one hop and awkward at three, because by then a traversal that follows every edge reaches most of the corpus, and narrowing it means having something to narrow on. A field on the edge is one way, which is what kind is here. Naming the edge tables more finely is the other: ->cites_background-> and ->cites_benchmark-> as separate relationships, so the hop picks the edges it wants and needs no filter at all. A field is the more flexible of the two, since it can be indexed, read back with the result, and combined at query time, while separate tables settle the question in the schema.
kind is the field that does most of the work. A background citation leads towards explanation, which is what a follow-up hop usually wants. A benchmark citation leads to numbers, and following those three deep is how a context window fills with nothing.
Putting kind on the edge rather than on the document is what lets the filter be pushed into the hop rather than applied to its output. Step 6 shows the difference that makes.
Topics are the second addition:
DEFINE TABLE OVERWRITE topic SCHEMAFULL;
DEFINE FIELD OVERWRITE name ON topic TYPE string;
DEFINE TABLE OVERWRITE about SCHEMALESS TYPE RELATION FROM document TO topic;A citation says "this document referenced that one". A topic says "these are about the same thing", which is a relationship no vector search recovers: two documents on one subject from different angles sit far apart in embedding space.
The chain that matters in the seed is four hops long, and every link is a background citation:
RELATE document:hnsw_tuning->cites->document:vec_index SET kind = 'background', year = 2025;
RELATE document:vec_index ->cites->document:storage SET kind = 'background', year = 2024;
RELATE document:storage ->cites->document:compaction SET kind = 'background', year = 2024;
RELATE document:compaction ->cites->document:disk_io SET kind = 'background', year = 2023;Load both files and open the shell:
surreal start --user root --pass secret
surreal import --endpoint http://localhost:8000 \
--user root --pass secret --ns ai --db graphrag schema.surql
surreal import --endpoint http://localhost:8000 \
--user root --pass secret --ns ai --db graphrag seed.surql
surreal sql --endpoint ws://localhost:8000 \
--user root --pass secret --ns ai --db graphrag --pretty < queries.surql
Step 2 - what similarity alone can offer
The question is "why does my vector index take so long to build?", which in the four toy axes [ databases, ml, web, devops ] is a vector weighted mostly towards databases, with a smaller amount on ml:
LET $q = [0.85, 0.40, 0.0, 0.15];
SELECT title, category,
math::round(vector::similarity::cosine(embedding, $q) * 1000) / 1000 AS sim
FROM document ORDER BY sim DESC LIMIT 7;[
{ title: 'Vector indexes in SurrealDB', category: 'database', sim: 0.998f },
{ title: 'Tuning HNSW parameters for recall', category: 'database', sim: 0.992f },
{ title: 'The storage engine: LSM trees', category: 'database', sim: 0.898f },
{ title: 'Compaction and write amplification', category: 'database', sim: 0.841f },
{ title: 'Embeddings for semantic search', category: 'ml', sim: 0.815f },
{ title: 'Chunking long documents', category: 'ml', sim: 0.802f },
{ title: 'Evaluating a RAG pipeline', category: 'ml', sim: 0.788f }
]A threshold has very little to work with here. These are cosine similarities, where 1.0 means the two vectors point the same way, so everything in that list looks like a strong match. "Compaction and write amplification" holds the answer and scores 0.841. "Chunking long documents" is irrelevant to the question and scores 0.802. Four hundredths separate the document you need from the document you do not, and nothing in that number says which is which.
That gap is not an artefact of toy vectors. It is what a similarity ranking is: a measure of resemblance, which is only a proxy for relevance, and a poor one once the answer stops resembling the question.
Take the two nearest as seeds and let the graph choose everything after that:
LET $seeds = (SELECT VALUE id FROM document WHERE embedding <|2, 40|> $q);
RETURN $seeds;[document:vec_index, document:hnsw_tuning]Step 3 - one hop, then two
This is lesson 04's hop, unchanged:
SELECT title, ->cites->document.title AS cites FROM $seeds;[
{ title: 'Vector indexes in SurrealDB', cites: ['The storage engine: LSM trees'] },
{ title: 'Tuning HNSW parameters for recall', cites: ['Vector indexes in SurrealDB'] }
]That is useful, but still short of the answer, so the next step is a second hop:
SELECT title, ->cites->document->cites->document.title AS hop2 FROM $seeds;[
{ title: 'Vector indexes in SurrealDB', hop2: ['Compaction and write amplification'] },
{ title: 'Tuning HNSW parameters for recall', hop2: ['The storage engine: LSM trees'] }
]There is the answer. And there is the problem with getting it this way: the depth is baked into the shape of the query. You had to know it was two by manually counting the number of ->cites->document occurrences, the two seeds needed different depths to reach the same document, and a corpus where the chain is sometimes four long would end up being mostly unreadable.
The second problem shows up as soon as the graph stops being a chain. Follow citations in both directions, which is what you want for context ("what does this cite" and "what cites this"), and count what comes back:
SELECT title,
(<->cites<->document<->cites<->document).len() AS paths,
(<->cites<->document<->cites<->document).distinct().len() AS nodes
FROM $seeds;[
{ title: 'Vector indexes in SurrealDB', paths: 32, nodes: 8 },
{ title: 'Tuning HNSW parameters for recall', paths: 8, nodes: 4 }
]A traversal returns every walk it made, rather than a de-duplicated list of what it reached. Thirty-two walks arrive at eight documents, so handing that straight to a prompt puts each document into it four times.
Step 4 - recursive traversal with a depth range
.{1..3+collect} walks the same relationship one to three times and returns the de-duplicated union of everything it reached:
RETURN $seeds.{1..3+collect}(->cites->document).group();[document:storage, document:vec_index, document:compaction, document:disk_io]One expression covers all three depths, and +collect returns each document once. The pieces:
.{1..3} | walk the relationship between one and three times |
+collect | return the union of every node reached, de-duplicated |
+path | return the walks instead of the destinations |
+shortest=<record> | return the shortest walk to one target |
(...) | the relationship to repeat, including any filters |
One constraint to know before you design around it: the depth in .{1..3} has to be a literal. It is part of the query, not a bound parameter, so a caller that varies depth varies the query text.
Step 5 - the frontier, and what depth 3 costs
Depth is the first parameter people set and the last one they measure. You can measure it with the query below:
RETURN [
{ depth: 1,
out_only: $seeds.{1..1+collect}(->cites->document).group().len(),
both_ways: $seeds.{1..1+collect}(<->cites<->document).group().len() },
{ depth: 2,
out_only: $seeds.{1..2+collect}(->cites->document).group().len(),
both_ways: $seeds.{1..2+collect}(<->cites<->document).group().len() },
{ depth: 3,
out_only: $seeds.{1..3+collect}(->cites->document).group().len(),
both_ways: $seeds.{1..3+collect}(<->cites<->document).group().len() }
];[
{ depth: 1, out_only: 2, both_ways: 4 },
{ depth: 2, out_only: 3, both_ways: 8 },
{ depth: 3, out_only: 4, both_ways: 12 }
]The corpus holds fourteen documents. Walking three hops in both directions reaches twelve of them:
RETURN $seeds.{1..3+collect}(<->cites<->document).group();[
document:hnsw_tuning, document:vec_index, document:embeddings, document:storage,
document:rag_eval, document:transformers, document:observability, document:compaction,
document:chunking, document:tokenisation, document:k8s, document:disk_io
]Kubernetes autoscaling is in there. So is tokenisation. Three hops in a well-connected citation graph is not "a bit more context", it is the corpus, sorted arbitrarily, and it will fill a context window with things the user did not ask about.
Direction is the first lever: ->cites-> alone reaches four documents where <->cites<-> reaches twelve. It is also the lever that quietly loses information, because "what cites this" is real context. The better lever is a filter on the edge itself, which the next step puts inside the hop.
Step 6 - push the filter into the hop
A condition inside the traversal expression applies at every hop rather than to the result:
LET $ctx = $seeds.{1..3+collect}(<->(cites WHERE kind = 'background')<->document).group();
RETURN $ctx;[
document:hnsw_tuning, document:vec_index, document:storage,
document:compaction, document:disk_io
]Twelve documents become five, and the five are the causal chain: index tuning, the index, the storage engine, compaction, disk IO. Both directions are still being followed, and the depth is still three. The only change is that the traversal refuses to leave the explanatory edges.
This is the difference between filtering a traversal and pruning one. A filter applied after the traversal has to expand the whole frontier first and then throw most of it away, and it cannot stop a walk that left the useful part of the graph at hop one from expanding for two more hops. A filter inside the hop never takes the step.
The parentheses are what make that true, and they are easy to lose. Square brackets accept a condition in the same position, but [WHERE ...] is the generic array filter, which walks the hop in full and filters the result afterwards. Both forms return the same documents here, so the only way to tell them apart is EXPLAIN, which puts a predicate on the GraphEdgeScan for the parenthesised form and not for the other. At three hops on a well-connected graph, that difference is most of the cost of the query.
The same syntax takes a filter on the node in the second set of parentheses, and both at once:
$seeds.{1..3+collect}(
<->(cites WHERE kind = 'background' AND year >= 2024)
<->(document WHERE category = 'database'))Our post Multi-hop graph traversal inside SurrealDB has the rest of the pruning toolkit, including per-hop frontier caps, what EXPLAIN says about edge indexes, and a trap to know about before you use [0..K] to cap a hop.
Step 7 - the graph decides membership, the vector decides order
The two rankings each do what they are good at:
SELECT title, category,
math::round(vector::similarity::cosine(embedding, $q) * 1000) / 1000 AS sim
FROM $ctx ORDER BY sim DESC;[
{ title: 'Vector indexes in SurrealDB', category: 'database', sim: 0.998f },
{ title: 'Tuning HNSW parameters for recall', category: 'database', sim: 0.992f },
{ title: 'The storage engine: LSM trees', category: 'database', sim: 0.898f },
{ title: 'Compaction and write amplification', category: 'database', sim: 0.841f },
{ title: 'Disk IO scheduling on NVMe', category: 'devops', sim: 0.607f }
]Step 2 ranked these differently. "Disk IO scheduling on NVMe" scores 0.607 and is in the context, four places below where a similarity search would ever have put it. "Chunking long documents" scores 0.802 but is not in the context at all, because nothing in the citation graph connects it to the question.
That inversion is the pattern. The graph decides what belongs in the context; the vector decides what order it goes in. Neither one can do the other's job: similarity has no idea what causes what, and a traversal has no idea what the question was.
Step 8 - return the walk, not just the destination
+path returns the walks, which is what you show a user who asks why:
RETURN document:hnsw_tuning.{1..3+path}(->(cites WHERE kind = 'background')->document)
.map(|$p| $p.map(|$d| $d.title).join(' -> '));['Vector indexes in SurrealDB -> The storage engine: LSM trees -> Compaction and write amplification']Outside this course, citing the documents behind an answer is now the ordinary expectation of a RAG system. An answer that shows the chain it followed to reach them is a different class of trustworthiness, and it costs one keyword. The walk starts at the first hop rather than at the seed.
When the question is about the connection itself, +shortest answers it directly:
RETURN document:hnsw_tuning.{..5+shortest=document:disk_io}(->cites->document)
.map(|$d| $d.title);[
'Vector indexes in SurrealDB',
'The storage engine: LSM trees',
'Compaction and write amplification',
'Disk IO scheduling on NVMe'
]That chain is the answer to a question nobody typed: "what does tuning an index have to do with disk scheduling?" is a path through the corpus rather than a similarity, which is why a vector search has no way to ask it.
Step 9 - the lateral hop
Citations only connect documents whose authors read each other. Two documents about the same subject, written from different angles, are often connected by nothing at all. Hop through the topic node to find them, and exclude what the citation traversal already collected:
LET $lateral = $seeds.{1..1+collect}(->about->topic<-about<-document)
.group()
.complement($ctx);
SELECT title, category,
math::round(vector::similarity::cosine(embedding, $q) * 1000) / 1000 AS sim
FROM $lateral ORDER BY sim DESC;[
{ title: 'Embeddings for semantic search', category: 'ml', sim: 0.815f },
{ title: 'Evaluating a RAG pipeline', category: 'ml', sim: 0.788f }
]Both of those documents are filed under "vector search", and neither the citation chain nor a top-3 vector search reaches either of them. ->about->topic<-about<-document is out to a hub and back, which is the cheapest lateral move in a graph and the easiest one to regret: a popular topic node is a hub, and one hop out and back through it can return most of the corpus. Keep the depth at one, and prune on the topic if the hub is large. Where a deeper lateral hop is genuinely needed, a TIMEOUT on the SELECT bounds what a bad case can cost, since it stops the query rather than the traversal.
Step 10 - the budget
A hop is fast, but the tokens it collects still have to fit the context window:
LET $ranked = (SELECT id, title, body.len() AS chars,
vector::similarity::cosine(embedding, $q) AS sim
FROM $ctx ORDER BY sim DESC LIMIT 4);
RETURN { docs: $ranked.title, chars: math::sum($ranked.chars) };{
docs: ['Vector indexes in SurrealDB', 'Tuning HNSW parameters for recall',
'The storage engine: LSM trees', 'Compaction and write amplification'],
chars: 475
}Four documents and 475 characters here, because the bodies are one sentence each. Substitute real documents and the same traversal at depth 3 returns tens of thousands of tokens, and the cut has to happen in the database rather than after the round trip.
The order of operations that works: seed with a vector search, expand with a pruned traversal, rank the union by similarity to the query, then cut the list with a LIMIT that fits the prompt. Ranking before cutting is what stops those places going to whichever document the traversal happened to reach first.
And to finish up the course, one final tip:
When not to go past one hop
Most questions do not need a second hop, and depth is not free. Three tests before adding one:
Does the answer live in a different document from the mention? A second hop helps when the corpus is causal or referential: citations, incident reports pointing at changes, tickets pointing at commits, entities pointing at documents. It adds nothing on a flat help centre where every article is self-contained.
Is there something to prune on? Step 5 is what an unpruned traversal does. If your edges carry no type, no weight and no date, add one before you add a hop, or the traversal is an expensive way to select the whole table.
Would you show the extra documents to a user? If a document three hops out would look like a non-sequitur in a citation list, it will read like one in a prompt too.
The way to settle it is to measure. Lesson 10 is the harness: run your golden set with the traversal and without it, and let recall@k (the share of the labelled answers each run retrieved) say whether depth 3 was worth the tokens. "Graph RAG" is a claim like any other.
Done the course!
Well done, you've reached the end of the course! Eleven lessons is only enough to get the shape of the subject, but you have now built the pieces that most retrieval systems are made of: vector search over an HNSW index, full-text search with BM25, a knowledge base with typed filters and a citation graph, two retrievers fused with Reciprocal Rank Fusion, a memory store that forgets and reinforces, an agent that writes its own queries against a prompt the database generates, three chunking strategies compared against each other, a golden set with recall@k and MRR to compare them with, and a traversal that reaches the document three hops away that no embedding would have found.
The last three lessons are the ones to come back to. Lessons 02 to 08 each add a mechanism, and it is easy to finish them believing the system works because every query returned something. Lesson 10 is the one that tells you whether it does.
Everything you ran needed one binary and one database. The vector index, the full-text index, the graph and the memory store were all features of the same engine, and an embedding model only entered where you chose to add one. That's easier to see with the queries in front of you than in the abstract.
Check out our other courses to keep going:
SurrealDB Fundamentals, where you'll build a simple ecommerce database, guided through video and text, and get a completion certificate at the end.
Schema internals and migrations, which goes deep on
DEFINEstatements, type safety and how to evolve a schema you already have in production.Aeon's Surreal Renaissance, where you'll immerse yourself in a fantasy book in which you follow the main character in a medieval future that uses SurrealDB to rebuild the world of our 21st century.
Tour of SurrealDB, the shortest of the courses, if you would like a quick pass over the whole feature set rather than a deep dive into one part of it.
And do drop by our community. The most active part is Discord - a good place to meet others building with SurrealDB, share what you're working on, get advice, and talk with the team. #showcase, #help and #general are good places to start, and if you build a retrieval layer out of any of this we'd like to see it.
Hope you enjoyed the course!
THE PLATFORM
Everything an application and its agents know. Five surfaces, one engine.
Database
Document, graph, vector, time-series and relational in one engine.

Agent Memory
What an agent learns, with its source and its time, in the same engine.

Cloud
Managed clusters in the regions you choose, scaled on demand.

Studio
Query, explore and design the schema from the browser.

MCP
Every model that speaks MCP reaches the database and the memory directly.

IN PRODUCTION
Trusted at scale. Samsung, Nvidia, Verizon, Tencent, and Walmart run on SurrealDB.
14,000+
Developers building on SurrealDB Cloud
4M+
Developers building on SurrealDB worldwide
FROM THE TEAMS
SurrealDB gives us a foundation where we can unify semantic search, knowledge graphs, and AI-driven decision making without stitching together multiple systems. Collapsing responsibility into SurrealDB has become our default engineering posture.
VP of Engineering, Later