
Many sources and agents, one context layer
This last lesson is a design one: there is nothing to install, and the SurrealQL in it sketches a system rather than something you run. The mechanisms underneath it are the ones lessons 04 to 07 already ran for real.
The question that needs seven systems
A customer has reported that something is broken, and a support engineer has picked up the case. They ask your assistant:
Northwind's checkout has been failing since Tuesday. What do we know?
While nothing about that question is exotic, no single system can answer it:
| Source | Shape | What it holds |
|---|---|---|
| Relational DB | rows, system of record | the customer, their plan, entitlements, which region they run in |
| Data lake | events, billions of rows | checkout attempts and error codes, per minute |
| Documents | unstructured | the checkout runbook, an architecture PDF nobody has opened since the rewrite |
| Jira | tickets | the open bug, its priority, who picked it up |
| GitHub | PRs and commits | the deploy that went out on Tuesday |
| Slack | threads | #checkout-eu, where someone already diagnosed it, and #security-incidents, where the same customer is being discussed for a very different reason |
| Notion | pages | the escalation policy and, three clicks away, a compensation review |
Because no single agent handles all of that well either, you split the work. An orchestrator reads the question and delegates: an account agent for who this customer is and what they're owed, an engineering agent for what changed and what broke. That's a good split: each agent has a smaller set of tools and a narrower brief.
Then comes the decision that matters. Splitting the reasoning across agents does not mean splitting the data access across sources, but that's what most people build, because that's how the connectors are sold.
Architecture A - glue in the agent
Each agent gets one tool per source, and the orchestrator concatenates whatever comes back:

It works and demos beautifully, but comes with a cost.
The aggregation point has no identity. Every connector authenticates as itself: a Slack bot token, a Jira service user, an integration installed workspace-wide, a lake role. Each of those credentials is deliberately broad, because it has to serve every user of the assistant. Each source then enforces its ACL correctly, against the wrong principal. Nowhere in the flow does anything evaluate what this human is allowed to see. The composition happens in the context window, and a context window has no access-control model to apply, because it is only a buffer of text.
Copied ACLs go stale. If you indexed the SaaS content to make retrieval fast, your permission snapshot is only as fresh as the last crawl. A channel made private on Tuesday is still public in your index until Thursday, and the record that leaks is the one somebody just decided to restrict, which is usually the record that mattered.
The leak is invisible. Each source logs an authorised read by its service account, so every audit trail is clean. Nothing logs the composition, the set of records this person was actually shown. There is no single place to ask "what did that user see?", either before you ship or after somebody notices.
Retrieval is per-source top-k, then concatenated. Every tool returns its own best five, and nothing ranks across them. The Slack thread that solves the case competes only against other Slack threads, while five weak matches from a chattier source take up the window. That's the problem lesson 05 solves with fusion, but fusion needs one ranking, and this architecture never has one.
Joins are inferred from strings. The word "Northwind" in a Slack message becomes the customer record by name-matching, in the model's head, at answer time. Internal codenames, two similarly-named accounts, a nickname somebody used in one thread: all become relationships the model asserts and cannot verify. Lesson 04's citation graph and lesson 06's graph hop exist because a traversed edge is a fact and an inferred join is a guess.
And it's slow. One round trip per source per sub-question, whole documents into the window, re-read on every hop: lesson 06's cold-start tax with seven times the surface area.
Now run it with the wrong person at the keyboard. A contractor on the integrations team asks the same question. The Slack bot was added to #security-incidents months ago so it could post alerts. The Notion integration was installed workspace-wide because scoping it was fiddly. The answer comes back fluent and useful, and folded into it is an unreleased vendor-breach disclosure naming this customer, plus a stray line from a compensation review that matched on a team name. There was no jailbreak here, no prompt injection and no malicious query. While every component did exactly what it was built to do, the trouble is that nowhere in this design was there a place to say no.
Prompt rules don't fix this. Lesson 07 makes the case on the write path: a system prompt is a request, and a model can be talked past it. The read path is worse, because there is nothing to talk past. The retrieval already succeeded, correctly, on behalf of a principal that was allowed. By the time the model sees the records, the decision has been made.
Architecture B - a context layer underneath
The agents and the sources are the same. One thing changes: the agents stop joining the data and start reasoning about it.

That one change has three effects, and they cover the list above.
One identity, evaluated before retrieval
The layer authenticates the human, using record access:
-- illustrative, not run in this lesson
DEFINE ACCESS staff ON DATABASE TYPE RECORD
SIGNIN (
SELECT * FROM staff
WHERE email = $email AND crypto::argon2::compare(pass, $pass)
)
DURATION FOR TOKEN 15m, FOR SESSION 12h;Signing in returns a token scoped to that person, and the agent carries it on every query it makes on their behalf, which is the one architectural rule that matters here: agents get the caller's token, never their own. $auth is then the caller's own staff record, with their groups, their team and whatever else you put on it, and permissions on the retrieval surface can read it directly:
-- illustrative, not run in this lesson
DEFINE TABLE context_item SCHEMAFULL PERMISSIONS
FOR select WHERE visibility = 'public'
OR $auth.id IN allowed_users
OR allowed_groups CONTAINSANY $auth.groups;What you get from that: the same query returns different records to different people, and the filter runs before ranking rather than after composition. Ask for the five nearest items to "Northwind checkout failing". An engineer who belongs to the security group gets the #checkout-eu thread and the vendor-breach disclosure. A contractor who belongs only to eng, running the identical query, gets the thread. Not a redacted version of the disclosure, not a refusal the model composed: the record is not in their result set at all.
| Item | Engineer (eng, security) | Contractor (eng) |
|---|---|---|
#checkout-eu thread (public) | returned | returned |
Vendor-breach disclosure (security only) | returned | not in the result set |
Field permissions give you the finer instrument, for the common case where the record is fine but one field isn't:
-- illustrative, not run in this lesson
DEFINE FIELD contract_value ON customer TYPE decimal
PERMISSIONS FOR select WHERE $auth.groups CONTAINS 'finance';The account agent can now tell a support engineer who the customer is and what they're entitled to, while the contract value comes back as NONE for anyone outside finance: one rule on one field, rather than two versions of a prompt and a hope.
Keep two ideas apart. The caller's identity fixes the scope: which records and fields exist as far as this session is concerned, and no agent in the fleet can exceed it, however it was prompted. What each agent may do within that scope is capability, and that's a separate decision: a distinct access method or database user per agent role, and tooling that only reaches the tables that role needs. The engineering agent working a checkout incident has no business reading contract values even when the human who asked could. The two compose as an intersection: the agent gets the smaller of what it's allowed and what its caller is allowed, so adding an agent to the fleet can never give anyone access they did not already have.
$authneeds record access. ADEFINE USER … ROLES VIEWER, lesson 07's guardrail, grants capability, not scope: it settles what kinds of statement the connection may run, and leaves$authempty. A root connection ignores table permissions altogether. The read path has to be bolted to an identity, or these clauses are decoration. Note also that a table with noPERMISSIONSclause returns nothing to a record user, so the layer is deny-by-default: you open tables up deliberately, one rule at a time.
One retrieval, ranked globally
With everything behind one permissioned surface, retrieval is a single query over everything the caller may see. That means the lesson 05 machinery finally applies across sources rather than within one: vector search over every item's embedding, BM25 over whatever text you actually hold, fused with RRF into one ranking, with lesson 04's relevance threshold keeping weak matches out of the window. There is one ordering to reason about and one token budget to spend, rather than seven of each.
If you carry provenance on every record (source, url, updated_at), the agent cites instead of asserting, and the engineer can click through to the Slack thread and see for themselves. If you rank first and fetch bodies only for the handful that survive, your token spend tracks the answer rather than the corpus.
One ontology, so entities join instead of matching
Retrieval hits one uniform surface: context_item, carrying a title, an embedding, provenance and the ACL fields. The facts live in typed entity tables (customer, service, deploy, incident) and the two are wired together by edges asserted once, at ingest, by the connector that actually knows the mapping:
slack_message → mentions → customer:northwind
jira_issue → about → service:checkout
pull_request → fixes → jira_issue:…
deploy → changed → service:checkoutThat is the ontology, and it is what makes "what do we know about Northwind's checkout?" a traversal rather than a string match. The join is stated once by something that can be held accountable for it, then followed a thousand times, instead of re-guessed by a language model on every question.
The other benefit is that you can check the answer. The chain from the thread, to the deploy that changed the service, to the ticket that fixes it, is a path in the graph. The agent's answer can be checked against it, which is a categorically different thing from a paragraph the model wrote confidently.
Where the data actually lives
None of this means "copy everything into SurrealDB". The pattern differs by source, and getting it wrong is how a context layer becomes a second data warehouse nobody trusts.
| Source | Pattern | What lands in SurrealDB |
|---|---|---|
| Relational DB (system of record) | CDC (change data capture) or periodic sync into typed tables | the entities and their relationships, the ontology's spine |
| Data lake | pre-aggregate at the source | rollups and a pointer, not events |
| Unstructured documents | materialise: chunk, embed, store | the chunks; here copying is worth the cost, because BM25 and vector search both need the text |
| SaaS (Slack, Notion, Jira, GitHub) | zero-copy reference | identity, title, URL, ACL, timestamps and an embedding, but not the body |
On the lake row: SurrealDB does not query external tables. There is no federation here. What works is running the aggregation where the events already are, on whatever schedule your question tolerates, and landing the result - checkout success rate per region per hour - as rows the agent can reason over, alongside a path for the one-off case where somebody has to go and look.
Zero-copy - SurrealDB as the ontology, not the copy
For the SaaS tools, the reference row holds everything you need to find the thing and nothing you need to read it: which system, the external ID, the URL, who may see it, when it last changed, and an embedding of the title or a short summary. The body stays in Slack.
The reason is that the SaaS tool is the system of record for the content and its ACL, and those two are the same fact. Copy the body and you have copied the ACL, which starts drifting the moment somebody changes a channel's membership. You also now owe deletion and retention in two places, which is the difference between "we index Slack" as an engineering decision and as a compliance conversation.
So the work splits in two. Ranking happens in SurrealDB, against permissions you control and can reason about, and the surviving handful of bodies are then fetched through the source API with the user's own token. Two live checks, one at each layer: yours before retrieval, theirs at fetch. If either says no, the body never reaches the window. Revocation stops being a re-index and becomes a record that no longer matches.
What you give up:
No BM25 over text you don't hold. Lesson 03's lexical recall only works on materialised text, so discovery leans on titles, summaries and the ontology. Titles are a weaker signal than chunks, and you will feel it on long threads.
A fetch on the critical path. Mitigated by only fetching the top few, and by a short-lived cache, but it is latency you didn't have when everything was local.
Source rate limits become your retrieval limits.
Materialising is the better call when the source is yours, when it's being retired, or when full-text recall over the body is the product. There is a middle path too, storing an LLM-written summary plus its embedding with no verbatim body. That gives better recall than a title, and it is still a copy as far as your compliance team is concerned.
What this buys, side by side
| Glue in the agent | Context layer | |
|---|---|---|
| Principal at read time | the connector's service account | the end user, via $auth |
| Where permissions are evaluated | in each source, per fetch, for the wrong principal | once, in the query, before ranking |
| ACL freshness | as stale as the last crawl | live at the source, or one row to update |
| Ranking | per-source top-k, concatenated | global, fused, thresholded |
| Cross-source joins | inferred from strings, per question | asserted once at ingest, then traversed |
| Round trips per question | one per source, per agent | one |
| Audit trail | N clean logs of service-account reads | one query, one identity, the record IDs returned |
| Dominant failure mode | over-retrieval that reads like a good answer | under-retrieval |
That last row deserves a second look, because it is the one that can keep a person up at night. If a context layer gets it wrong, somebody does not get a record they needed and can see that they didn't, which makes it a bug you can measure, reproduce and fix. If the glue architecture gets it wrong, somebody gets a record nobody meant them to see, phrased confidently and attributed to nothing, while every log in the building shows green.
Adopting it incrementally
Start with one source and add the rest as questions need them.
- 1.
Build the spine. Sync the entities and relationships from your relational system of record. That alone gives the agents somewhere to resolve "Northwind" to a record instead of a string.
- 2.
Add one SaaS source as references. Pick the noisiest one, usually Slack, and land identity, ACL and an embedding. No bodies.
- 3.
Put that source's retrieval behind one permissioned query, authenticated as the human. Delete the tool that fetched it with a bot token.
Then keep adding sources. Steps 1-3 don't get redone: the permission rule and the retrieval query don't change shape as the fourth and fifth source arrive. That's why the security property holds as the system grows, rather than degrading with every integration somebody adds in a hurry.
What to take from this lesson
If you take one thing away, make it this. Decide where in your system the identity of the person asking is known, and put retrieval there. Everything else in this lesson follows from that decision, and in Architecture A there is no such place to point at.
The next three lessons go back to the retrieval layer itself with the same question in mind: whether what you built actually returns the right documents. Lesson 09 asks how big a chunk should be, lesson 10 builds the harness that answers questions like that with a number, and lesson 11 follows the citations when the document that explains a problem is not the document that mentions it.
THE PLATFORM
Everything an application and its agents know. Five surfaces, one engine.
Database
Document, graph, vector, time-series and relational in one engine.

Agent Memory
What an agent learns, with its source and its time, in the same engine.

Cloud
Managed clusters in the regions you choose, scaled on demand.

Studio
Query, explore and design the schema from the browser.

MCP
Every model that speaks MCP reaches the database and the memory directly.

IN PRODUCTION
Trusted at scale. Samsung, Nvidia, Verizon, Tencent, and Walmart run on SurrealDB.
14,000+
Developers building on SurrealDB Cloud
4M+
Developers building on SurrealDB worldwide
FROM THE TEAMS
SurrealDB gives us a foundation where we can unify semantic search, knowledge graphs, and AI-driven decision making without stitching together multiple systems. Collapsing responsibility into SurrealDB has become our default engineering posture.
VP of Engineering, Later