Skip to content
Course content preview

SurrealDB for AI Engineers

SurrealDB for AI Engineers is an eleven-lesson course that takes you from "I've called an LLM API" to an agent that retrieves by meaning, remembers what it learned, and writes its own queries. Lesson 08 weighs two architectures for the case where the answer lives across several systems and more than one agent is asking. The last three lessons turn back on the retrieval layer itself and ask whether it returns the right documents, which is the question the earlier lessons give you no way to answer.

This course is for anyone who writes software, has called an LLM API, and is now looking at the retrieval layer underneath it. Prior SurrealDB experience isn't needed, since the track introduces every feature where it first comes up. An embedding model and a cloud account aren't needed either: the examples use tiny hand-written vectors so that the maths stays readable, and every lesson that uses them ends by showing what to change to swap in a real model.

The one thing you will want is the surreal binary, which you can download via a single command:

curl -sSf https://install.surrealdb.com | sh
surreal version

As one of SurrealDB University's modular courses, this one also focuses on a single topic and is just a few chapters in length. It runs over eleven pages. Lesson 01 is background reading with no SurrealDB in it, so you can skip it if you already build with LLMs. Lessons 02 and 03 are the retrieval primitives, one idea each with a small runnable example, and every later lesson builds on one or both. Lessons 04 to 07 are the applied patterns. Lesson 08 is design rather than code, with nothing to install. Lessons 09 to 11 are about retrieval quality: how you split a document, how you measure whether the split helped, and what to do when the answer is a hop away from anything the question resembles.

From lesson 02 onwards you start a local surreal server and load small .surql files. Each hands-on lesson gives you a schema, a seed, and a query file to save and run. Keep the server in one terminal and run the import and surreal sql commands in another.

Lesson

What you get

01

AI foundations

How LLMs work (tokens, context windows, why they hallucinate) and the difference between prompt engineering and context engineering.

02

Vector embeddings and vector search

What an embedding is, exact search with no index, DEFINE INDEX … HNSW and every parameter, the KNN operator, vector::distance::knn(), and how to choose an embedding model.

03

Full-text search and BM25

What lexical search scores, DEFINE ANALYZER and stemming, DEFINE INDEX … FULLTEXT … BM25(k1, b), the @@ operator, search::score() and search::highlight().

04

RAG knowledge base

Semantic retrieval with typed filters and a citation graph, in one query. Relevance thresholds that keep weak matches out of the prompt, and the surviving documents rendered into it.

05

Hybrid search and reranking

BM25 and vector search over the same table, fused with Reciprocal Rank Fusion. Hand-rolled first, then with the built-in search::rrf.

06

Agent memory store

Recall that blends similarity, recency and importance; graph-assembled context; forgetting via expires_at; and a write-back that makes a memory count for more each time it is recalled.

07

Text-to-SurrealQL

A prompt generated from the database's own DEFINE statements, so it can't drift from the schema. Dynamic few-shot examples, the three ways generation fails, and guardrails the model can't talk past.

08

Many sources, many agents, one context layer

Two architectures for a multi-agent system over a data lake, a relational DB, documents and SaaS. Why the agent-as-aggregator hands people records they can't see, and what record access, $auth and zero-copy references change.

09

Chunking strategies

Three ways to split a document compared on the same corpus, why a chunk that covers more ground points nowhere, and the record link that lets you retrieve something small and generate from something large.

10

Evaluating retrieval quality

A golden set with the answers labelled, recall@k, precision@k and MRR as SurrealQL functions, and and the harness that settles which retriever is better.

11

Graph RAG beyond one hop

Multi-hop traversal for when the document that explains a problem is not the document that mentions it. Depth limits, pruning, +shortest, and how to tell whether the extra hops are worth their tokens.

               01  AI foundations
                        │
           ┌────────────┴─────────────┐
 02  vector embeddings     03  full-text and BM25
           │                          │
04  RAG knowledge base                │
           └────────────┬─────────────┘
                        │
                05  hybrid search
                        │
                06  agent memory
                        │
              07  text-to-SurrealQL
                        │
                08  context layer
                        │
           ┌────────────┴─────────────┐
 09  chunking strategies   11  graph RAG, multi-hop
           │                          │
           └────────────┬─────────────┘
                        │
            10  evaluating retrieval

Lesson 04 uses the vector search from lesson 02, and it's first because it's the simpler of the two applied lessons. Lesson 05 is where vector search and full-text search meet. Lesson 06 adds memory on top of retrieval, and lesson 07 lets the agent write its own queries. Lesson 08 is about how you put the pieces together once the data lives in several systems.

Lessons 09 and 11 are the two ways of changing what gets retrieved: cut the documents up differently, or follow the links between them. Lesson 10 sits under both, because a golden set - a list of questions paired with the documents that answer them, labelled by hand - is how you find out which of them helped. It is drawn last because it is easier to write once you have two retrievers to compare, though you can read it directly after lesson 09 if you would rather measure before you add anything else.

Each lesson stands alone if you would rather jump straight in, and the cross-references will tell you what parts you are skipping when you do.

While you're here, feel free to set up an account with us, browse the full docs or join the largest community of SurrealDB users in one place.

THE PLATFORM

Everything an application and its agents know. Five surfaces, one engine.

IN PRODUCTION

Trusted at scale. Samsung, Nvidia, Verizon, Tencent, and Walmart run on SurrealDB.

14,000+

Developers building on SurrealDB Cloud

4M+

Developers building on SurrealDB worldwide

FROM THE TEAMS

SurrealDB gives us a foundation where we can unify semantic search, knowledge graphs, and AI-driven decision making without stitching together multiple systems. Collapsing responsibility into SurrealDB has become our default engineering posture.
Justin Foley

VP of Engineering, Later

SurrealDB

The context and memory layer for AI agents

Database. Graphs, vectors, documents and relational data in one engine, in a single ACID transaction.
Agent Memory. Connects and retrieves context wherever your data lives, every fact carrying its source.
Cloud. Fully managed, in the cloud provider and region you choose.

Explore with AI

Copyright © 2026 SurrealDB Ltd. Registered in England and Wales. Company no. 13615201

Registered address: 3rd Floor 1 Ashley Road, Altrincham, Cheshire, WA14 2DT, United Kingdom

Trading address: Huckletree Oxford Circus, 213 Oxford Street, London, W1D 2LG, United Kingdom