Skip to content
Course content preview

Many sources and agents, one context layer

This last lesson is a design one: there is nothing to install, and the SurrealQL in it sketches a system rather than something you run. The mechanisms underneath it are the ones lessons 04 to 07 already ran for real.

A customer has reported that something is broken, and a support engineer has picked up the case. They ask your assistant:

Northwind's checkout has been failing since Tuesday. What do we know?

While nothing about that question is exotic, no single system can answer it:

SourceShapeWhat it holds
Relational DBrows, system of recordthe customer, their plan, entitlements, which region they run in
Data lakeevents, billions of rowscheckout attempts and error codes, per minute
Documentsunstructuredthe checkout runbook, an architecture PDF nobody has opened since the rewrite
Jiraticketsthe open bug, its priority, who picked it up
GitHubPRs and commitsthe deploy that went out on Tuesday
Slackthreads#checkout-eu, where someone already diagnosed it, and #security-incidents, where the same customer is being discussed for a very different reason
Notionpagesthe escalation policy and, three clicks away, a compensation review

Because no single agent handles all of that well either, you split the work. An orchestrator reads the question and delegates: an account agent for who this customer is and what they're owed, an engineering agent for what changed and what broke. That's a good split: each agent has a smaller set of tools and a narrower brief.

Then comes the decision that matters. Splitting the reasoning across agents does not mean splitting the data access across sources, but that's what most people build, because that's how the connectors are sold.

Each agent gets one tool per source, and the orchestrator concatenates whatever comes back:

Architecture A: each agent holds one tool per source, and the orchestrator concatenates whatever comes back

It works and demos beautifully, but comes with a cost.

The aggregation point has no identity. Every connector authenticates as itself: a Slack bot token, a Jira service user, an integration installed workspace-wide, a lake role. Each of those credentials is deliberately broad, because it has to serve every user of the assistant. Each source then enforces its ACL correctly, against the wrong principal. Nowhere in the flow does anything evaluate what this human is allowed to see. The composition happens in the context window, and a context window has no access-control model to apply, because it is only a buffer of text.

Copied ACLs go stale. If you indexed the SaaS content to make retrieval fast, your permission snapshot is only as fresh as the last crawl. A channel made private on Tuesday is still public in your index until Thursday, and the record that leaks is the one somebody just decided to restrict, which is usually the record that mattered.

The leak is invisible. Each source logs an authorised read by its service account, so every audit trail is clean. Nothing logs the composition, the set of records this person was actually shown. There is no single place to ask "what did that user see?", either before you ship or after somebody notices.

Retrieval is per-source top-k, then concatenated. Every tool returns its own best five, and nothing ranks across them. The Slack thread that solves the case competes only against other Slack threads, while five weak matches from a chattier source take up the window. That's the problem lesson 05 solves with fusion, but fusion needs one ranking, and this architecture never has one.

Joins are inferred from strings. The word "Northwind" in a Slack message becomes the customer record by name-matching, in the model's head, at answer time. Internal codenames, two similarly-named accounts, a nickname somebody used in one thread: all become relationships the model asserts and cannot verify. Lesson 04's citation graph and lesson 06's graph hop exist because a traversed edge is a fact and an inferred join is a guess.

And it's slow. One round trip per source per sub-question, whole documents into the window, re-read on every hop: lesson 06's cold-start tax with seven times the surface area.

Now run it with the wrong person at the keyboard. A contractor on the integrations team asks the same question. The Slack bot was added to #security-incidents months ago so it could post alerts. The Notion integration was installed workspace-wide because scoping it was fiddly. The answer comes back fluent and useful, and folded into it is an unreleased vendor-breach disclosure naming this customer, plus a stray line from a compensation review that matched on a team name. There was no jailbreak here, no prompt injection and no malicious query. While every component did exactly what it was built to do, the trouble is that nowhere in this design was there a place to say no.

Prompt rules don't fix this. Lesson 07 makes the case on the write path: a system prompt is a request, and a model can be talked past it. The read path is worse, because there is nothing to talk past. The retrieval already succeeded, correctly, on behalf of a principal that was allowed. By the time the model sees the records, the decision has been made.

The agents and the sources are the same. One thing changes: the agents stop joining the data and start reasoning about it.

Architecture B: the agents reason, and one context layer underneath them joins the data under a single identity

That one change has three effects, and they cover the list above.

The layer authenticates the human, using record access:

-- illustrative, not run in this lesson
DEFINE ACCESS staff ON DATABASE TYPE RECORD
    SIGNIN (
        SELECT * FROM staff
        WHERE email = $email AND crypto::argon2::compare(pass, $pass)
    )
    DURATION FOR TOKEN 15m, FOR SESSION 12h;

Signing in returns a token scoped to that person, and the agent carries it on every query it makes on their behalf, which is the one architectural rule that matters here: agents get the caller's token, never their own. $auth is then the caller's own staff record, with their groups, their team and whatever else you put on it, and permissions on the retrieval surface can read it directly:

-- illustrative, not run in this lesson
DEFINE TABLE context_item SCHEMAFULL PERMISSIONS
    FOR select WHERE visibility = 'public'
        OR $auth.id IN allowed_users
        OR allowed_groups CONTAINSANY $auth.groups;

What you get from that: the same query returns different records to different people, and the filter runs before ranking rather than after composition. Ask for the five nearest items to "Northwind checkout failing". An engineer who belongs to the security group gets the #checkout-eu thread and the vendor-breach disclosure. A contractor who belongs only to eng, running the identical query, gets the thread. Not a redacted version of the disclosure, not a refusal the model composed: the record is not in their result set at all.

ItemEngineer (eng, security)Contractor (eng)
#checkout-eu thread (public)returnedreturned
Vendor-breach disclosure (security only)returnednot in the result set

Field permissions give you the finer instrument, for the common case where the record is fine but one field isn't:

-- illustrative, not run in this lesson
DEFINE FIELD contract_value ON customer TYPE decimal
    PERMISSIONS FOR select WHERE $auth.groups CONTAINS 'finance';

The account agent can now tell a support engineer who the customer is and what they're entitled to, while the contract value comes back as NONE for anyone outside finance: one rule on one field, rather than two versions of a prompt and a hope.

Keep two ideas apart. The caller's identity fixes the scope: which records and fields exist as far as this session is concerned, and no agent in the fleet can exceed it, however it was prompted. What each agent may do within that scope is capability, and that's a separate decision: a distinct access method or database user per agent role, and tooling that only reaches the tables that role needs. The engineering agent working a checkout incident has no business reading contract values even when the human who asked could. The two compose as an intersection: the agent gets the smaller of what it's allowed and what its caller is allowed, so adding an agent to the fleet can never give anyone access they did not already have.

$auth needs record access. A DEFINE USER … ROLES VIEWER, lesson 07's guardrail, grants capability, not scope: it settles what kinds of statement the connection may run, and leaves $auth empty. A root connection ignores table permissions altogether. The read path has to be bolted to an identity, or these clauses are decoration. Note also that a table with no PERMISSIONS clause returns nothing to a record user, so the layer is deny-by-default: you open tables up deliberately, one rule at a time.

With everything behind one permissioned surface, retrieval is a single query over everything the caller may see. That means the lesson 05 machinery finally applies across sources rather than within one: vector search over every item's embedding, BM25 over whatever text you actually hold, fused with RRF into one ranking, with lesson 04's relevance threshold keeping weak matches out of the window. There is one ordering to reason about and one token budget to spend, rather than seven of each.

If you carry provenance on every record (source, url, updated_at), the agent cites instead of asserting, and the engineer can click through to the Slack thread and see for themselves. If you rank first and fetch bodies only for the handful that survive, your token spend tracks the answer rather than the corpus.

Retrieval hits one uniform surface: context_item, carrying a title, an embedding, provenance and the ACL fields. The facts live in typed entity tables (customer, service, deploy, incident) and the two are wired together by edges asserted once, at ingest, by the connector that actually knows the mapping:

slack_message → mentions → customer:northwind
jira_issue    → about    → service:checkout
pull_request  → fixes    → jira_issue:…
deploy        → changed  → service:checkout

That is the ontology, and it is what makes "what do we know about Northwind's checkout?" a traversal rather than a string match. The join is stated once by something that can be held accountable for it, then followed a thousand times, instead of re-guessed by a language model on every question.

The other benefit is that you can check the answer. The chain from the thread, to the deploy that changed the service, to the ticket that fixes it, is a path in the graph. The agent's answer can be checked against it, which is a categorically different thing from a paragraph the model wrote confidently.

None of this means "copy everything into SurrealDB". The pattern differs by source, and getting it wrong is how a context layer becomes a second data warehouse nobody trusts.

SourcePatternWhat lands in SurrealDB
Relational DB (system of record)CDC (change data capture) or periodic sync into typed tablesthe entities and their relationships, the ontology's spine
Data lakepre-aggregate at the sourcerollups and a pointer, not events
Unstructured documentsmaterialise: chunk, embed, storethe chunks; here copying is worth the cost, because BM25 and vector search both need the text
SaaS (Slack, Notion, Jira, GitHub)zero-copy referenceidentity, title, URL, ACL, timestamps and an embedding, but not the body

On the lake row: SurrealDB does not query external tables. There is no federation here. What works is running the aggregation where the events already are, on whatever schedule your question tolerates, and landing the result - checkout success rate per region per hour - as rows the agent can reason over, alongside a path for the one-off case where somebody has to go and look.

For the SaaS tools, the reference row holds everything you need to find the thing and nothing you need to read it: which system, the external ID, the URL, who may see it, when it last changed, and an embedding of the title or a short summary. The body stays in Slack.

The reason is that the SaaS tool is the system of record for the content and its ACL, and those two are the same fact. Copy the body and you have copied the ACL, which starts drifting the moment somebody changes a channel's membership. You also now owe deletion and retention in two places, which is the difference between "we index Slack" as an engineering decision and as a compliance conversation.

So the work splits in two. Ranking happens in SurrealDB, against permissions you control and can reason about, and the surviving handful of bodies are then fetched through the source API with the user's own token. Two live checks, one at each layer: yours before retrieval, theirs at fetch. If either says no, the body never reaches the window. Revocation stops being a re-index and becomes a record that no longer matches.

What you give up:

  • No BM25 over text you don't hold. Lesson 03's lexical recall only works on materialised text, so discovery leans on titles, summaries and the ontology. Titles are a weaker signal than chunks, and you will feel it on long threads.

  • A fetch on the critical path. Mitigated by only fetching the top few, and by a short-lived cache, but it is latency you didn't have when everything was local.

  • Source rate limits become your retrieval limits.

Materialising is the better call when the source is yours, when it's being retired, or when full-text recall over the body is the product. There is a middle path too, storing an LLM-written summary plus its embedding with no verbatim body. That gives better recall than a title, and it is still a copy as far as your compliance team is concerned.

Glue in the agentContext layer
Principal at read timethe connector's service accountthe end user, via $auth
Where permissions are evaluatedin each source, per fetch, for the wrong principalonce, in the query, before ranking
ACL freshnessas stale as the last crawllive at the source, or one row to update
Rankingper-source top-k, concatenatedglobal, fused, thresholded
Cross-source joinsinferred from strings, per questionasserted once at ingest, then traversed
Round trips per questionone per source, per agentone
Audit trailN clean logs of service-account readsone query, one identity, the record IDs returned
Dominant failure modeover-retrieval that reads like a good answerunder-retrieval

That last row deserves a second look, because it is the one that can keep a person up at night. If a context layer gets it wrong, somebody does not get a record they needed and can see that they didn't, which makes it a bug you can measure, reproduce and fix. If the glue architecture gets it wrong, somebody gets a record nobody meant them to see, phrased confidently and attributed to nothing, while every log in the building shows green.

Start with one source and add the rest as questions need them.

  1. 1.

    Build the spine. Sync the entities and relationships from your relational system of record. That alone gives the agents somewhere to resolve "Northwind" to a record instead of a string.

  2. 2.

    Add one SaaS source as references. Pick the noisiest one, usually Slack, and land identity, ACL and an embedding. No bodies.

  3. 3.

    Put that source's retrieval behind one permissioned query, authenticated as the human. Delete the tool that fetched it with a bot token.

Then keep adding sources. Steps 1-3 don't get redone: the permission rule and the retrieval query don't change shape as the fourth and fifth source arrive. That's why the security property holds as the system grows, rather than degrading with every integration somebody adds in a hurry.

If you take one thing away, make it this. Decide where in your system the identity of the person asking is known, and put retrieval there. Everything else in this lesson follows from that decision, and in Architecture A there is no such place to point at.

The next three lessons go back to the retrieval layer itself with the same question in mind: whether what you built actually returns the right documents. Lesson 09 asks how big a chunk should be, lesson 10 builds the harness that answers questions like that with a number, and lesson 11 follows the citations when the document that explains a problem is not the document that mentions it.

THE PLATFORM

Everything an application and its agents know. Five surfaces, one engine.

IN PRODUCTION

Trusted at scale. Samsung, Nvidia, Verizon, Tencent, and Walmart run on SurrealDB.

14,000+

Developers building on SurrealDB Cloud

4M+

Developers building on SurrealDB worldwide

FROM THE TEAMS

SurrealDB gives us a foundation where we can unify semantic search, knowledge graphs, and AI-driven decision making without stitching together multiple systems. Collapsing responsibility into SurrealDB has become our default engineering posture.
Justin Foley

VP of Engineering, Later

SurrealDB

The context and memory layer for AI agents

Database. Graphs, vectors, documents and relational data in one engine, in a single ACID transaction.
Agent Memory. Connects and retrieves context wherever your data lives, every fact carrying its source.
Cloud. Fully managed, in the cloud provider and region you choose.

Explore with AI

Copyright © 2026 SurrealDB Ltd. Registered in England and Wales. Company no. 13615201

Registered address: 3rd Floor 1 Ashley Road, Altrincham, Cheshire, WA14 2DT, United Kingdom

Trading address: Huckletree Oxford Circus, 213 Oxford Street, London, W1D 2LG, United Kingdom