
SurrealDB for AI Engineers
SurrealDB for AI Engineers is an eleven-lesson course that takes you from "I've called an LLM API" to an agent that retrieves by meaning, remembers what it learned, and writes its own queries. Lesson 08 weighs two architectures for the case where the answer lives across several systems and more than one agent is asking. The last three lessons turn back on the retrieval layer itself and ask whether it returns the right documents, which is the question the earlier lessons give you no way to answer.
Who this is for
This course is for anyone who writes software, has called an LLM API, and is now looking at the retrieval layer underneath it. Prior SurrealDB experience isn't needed, since the track introduces every feature where it first comes up. An embedding model and a cloud account aren't needed either: the examples use tiny hand-written vectors so that the maths stays readable, and every lesson that uses them ends by showing what to change to swap in a real model.
The one thing you will want is the surreal binary, which you can download via a single command:
curl -sSf https://install.surrealdb.com | sh
surreal versionThe lessons
As one of SurrealDB University's modular courses, this one also focuses on a single topic and is just a few chapters in length. It runs over eleven pages. Lesson 01 is background reading with no SurrealDB in it, so you can skip it if you already build with LLMs. Lessons 02 and 03 are the retrieval primitives, one idea each with a small runnable example, and every later lesson builds on one or both. Lessons 04 to 07 are the applied patterns. Lesson 08 is design rather than code, with nothing to install. Lessons 09 to 11 are about retrieval quality: how you split a document, how you measure whether the split helped, and what to do when the answer is a hop away from anything the question resembles.
From lesson 02 onwards you start a local surreal server and load small .surql files. Each hands-on lesson gives you a schema, a seed, and a query file to save and run. Keep the server in one terminal and run the import and surreal sql commands in another.
Lesson | What you get | |
|---|---|---|
01 | How LLMs work (tokens, context windows, why they hallucinate) and the difference between prompt engineering and context engineering. | |
02 | What an embedding is, exact search with no index, | |
03 | What lexical search scores, | |
04 | Semantic retrieval with typed filters and a citation graph, in one query. Relevance thresholds that keep weak matches out of the prompt, and the surviving documents rendered into it. | |
05 | BM25 and vector search over the same table, fused with Reciprocal Rank Fusion. Hand-rolled first, then with the built-in | |
06 | Recall that blends similarity, recency and importance; graph-assembled context; forgetting via | |
07 | A prompt generated from the database's own | |
08 | Two architectures for a multi-agent system over a data lake, a relational DB, documents and SaaS. Why the agent-as-aggregator hands people records they can't see, and what record access, | |
09 | Three ways to split a document compared on the same corpus, why a chunk that covers more ground points nowhere, and the record link that lets you retrieve something small and generate from something large. | |
10 | A golden set with the answers labelled, recall@k, precision@k and MRR as SurrealQL functions, and and the harness that settles which retriever is better. | |
11 | Multi-hop traversal for when the document that explains a problem is not the document that mentions it. Depth limits, pruning, |
How the lessons connect
01 AI foundations
│
┌────────────┴─────────────┐
02 vector embeddings 03 full-text and BM25
│ │
04 RAG knowledge base │
└────────────┬─────────────┘
│
05 hybrid search
│
06 agent memory
│
07 text-to-SurrealQL
│
08 context layer
│
┌────────────┴─────────────┐
09 chunking strategies 11 graph RAG, multi-hop
│ │
└────────────┬─────────────┘
│
10 evaluating retrievalLesson 04 uses the vector search from lesson 02, and it's first because it's the simpler of the two applied lessons. Lesson 05 is where vector search and full-text search meet. Lesson 06 adds memory on top of retrieval, and lesson 07 lets the agent write its own queries. Lesson 08 is about how you put the pieces together once the data lives in several systems.
Lessons 09 and 11 are the two ways of changing what gets retrieved: cut the documents up differently, or follow the links between them. Lesson 10 sits under both, because a golden set - a list of questions paired with the documents that answer them, labelled by hand - is how you find out which of them helped. It is drawn last because it is easier to write once you have two retrievers to compare, though you can read it directly after lesson 09 if you would rather measure before you add anything else.
Each lesson stands alone if you would rather jump straight in, and the cross-references will tell you what parts you are skipping when you do.
Related links
While you're here, feel free to set up an account with us, browse the full docs or join the largest community of SurrealDB users in one place.
THE PLATFORM
Everything an application and its agents know. Five surfaces, one engine.
Database
Document, graph, vector, time-series and relational in one engine.

Agent Memory
What an agent learns, with its source and its time, in the same engine.

Cloud
Managed clusters in the regions you choose, scaled on demand.

Studio
Query, explore and design the schema from the browser.

MCP
Every model that speaks MCP reaches the database and the memory directly.

IN PRODUCTION
Trusted at scale. Samsung, Nvidia, Verizon, Tencent, and Walmart run on SurrealDB.
14,000+
Developers building on SurrealDB Cloud
4M+
Developers building on SurrealDB worldwide
FROM THE TEAMS
SurrealDB gives us a foundation where we can unify semantic search, knowledge graphs, and AI-driven decision making without stitching together multiple systems. Collapsing responsibility into SurrealDB has become our default engineering posture.
VP of Engineering, Later