
AI foundations
This first lesson is the background the rest of the course uses: how a language model works, and the difference between writing a prompt and choosing what goes into the window. The reading is all external. Lessons 02 to 08 assume you've seen these ideas.
The later lessons use the usual terms - context window, context rot, attention budget, the agent loop - and put a database underneath them. Each link below is here because a later lesson needs it, and the note says which one.
How LLMs work
A language model does one thing: given a sequence of tokens, predict the next one.
Everything else (answering questions, writing code, holding a conversation) is that operation in a loop, each prediction fed back in.
Two things follow from that, and both show up for the rest of the course:
- 1.
The context window is the model's entire world at inference time: no memory between calls, no access to your data, so a fact that isn't in the window doesn't exist.
- 2.
Nothing in predict a plausible next token checks whether the continuation is true - hallucination is what the mechanism does when the window doesn't hold the answer. The fix is to put the answer there.
Resource | What it covers |
|---|---|
Transformers, the tech behind LLMs 3Blue1Brown, 27 minutes | Grant Sanderson builds a transformer up one piece at a time, animating what an embedding, an attention head and a softmax are each doing to the numbers. If you watch only one thing here, watch this one: lesson 02's talk of vectors and closeness in meaning gets a lot less abstract once you have seen the geometry move. |
Intro to Large Language Models Andrej Karpathy, one hour | Pretraining, fine-tuning and where capabilities actually come from, at a brisk pace. Karpathy is blunt about what models can't do, which most intros skip. He also draws a line this course sticks to: capability comes from training, while knowledge of your data has to arrive at inference time. That's why nothing from lesson 02 onwards fine-tunes a model. |
Jay Alammar | The diagram set that everyone else borrows from. It covers similar ground to the 3Blue1Brown video but on the page, which is easier when you want to stop on one step and stare at it. Its walk through how a token becomes a vector is the piece lesson 02 builds on most directly. |
Tat Dat Duong | A tool rather than a reading. Paste in any text and watch it break into tokens. Two minutes here makes the context window concrete. Notice how little a bare error code or an order number gives a model to work with: those are the queries lesson 03 is for. |

Four lines, twenty characters each, run through gpt-4o. The token count climbs with unfamiliarity rather than with length: I like cats and dogs costs 5, the same sentence with SurrealDB in it costs 6, an error code and an order number cost 7, and an invented word costs 8. Every step down that list leaves a model less to work with, which is the blind spot lesson 03 picks up. Screenshot of Tiktokenizer, built by Tat Dat Duong and MIT licensed.
Prompt engineering and context engineering
Prompt engineering is the wording of the instruction: role, format, examples. Write it once and reuse it.
Context engineering is deciding what information occupies the finite window on each request, and where it comes from. You solve that at runtime, on every call, against data that changes. It's a retrieval problem, which makes it a database problem.
Lessons 04 to 07 are all context engineering: retrieving the right documents, fusing two rankings, recalling what the agent learned last week, and generating the prompt itself out of live schema so it can't go stale.
When the data keeps changing. A store that only ever grows ends up holding the old version of a fact next to the new one, and a retriever with no idea of time will hand a model both. SurrealDB's Agent Memory is built around that: it keeps what was said, what is true now, and what used to be true as three separate things, so a superseded fact can be closed rather than left to compete with its replacement. Lesson 06 builds a small version of the same idea by hand, with an expiry field and a write-back on recall.
Resource | What it covers |
|---|---|
Effective context engineering for AI agents Anthropic, September 2025 | The distinction set out in full by the people who popularised the term. The context rot section is the useful bit: past a certain size, adding tokens to the window measurably degrades the model's recall, so you want less in the window, not more. It also covers compaction, structured note-taking and sub-agent architectures as three ways of living within an attention budget. Lesson 04 turns that into a |
Anthropic | The practical checklist: be explicit, give examples, let the model think, assign it a role. Skim it once now and come back to it when a prompt misbehaves. It works better as a reference than as a read. Lesson 07 uses this checklist almost line by line, except the database assembles the prompt rather than a person. |
Anthropic, December 2024 | Where retrieval sits inside an agent loop, and the argument for simple composable patterns over frameworks. The post now opens with a warning that its tooling advice has dated; the patterns have held up rather better than the tools around them. The agent loop it describes is the one lesson 06 writes back into, and the multi-agent shape it sketches is what lesson 08 takes apart. |
Lilian Weng, about half an hour's read | A survey, organised as planning, memory and tool use. The memory section is the bit lesson 06 follows most closely, and its detour through maximum inner product search is the same retrieval problem that lesson 02 hands to an HNSW index. |
Where this goes next
The window is finite, and nobody gets to choose what the user asks. So the question is: given a question you never anticipated, what goes in the window?
There are two ways to find that content, by meaning and by literal term. Lesson 02 covers the first, lesson 03 the second, and every later lesson leans on one or both.
THE PLATFORM
Everything an application and its agents know. Five surfaces, one engine.
Database
Document, graph, vector, time-series and relational in one engine.

Agent Memory
What an agent learns, with its source and its time, in the same engine.

Cloud
Managed clusters in the regions you choose, scaled on demand.

Studio
Query, explore and design the schema from the browser.

MCP
Every model that speaks MCP reaches the database and the memory directly.

IN PRODUCTION
Trusted at scale. Samsung, Nvidia, Verizon, Tencent, and Walmart run on SurrealDB.
14,000+
Developers building on SurrealDB Cloud
4M+
Developers building on SurrealDB worldwide
FROM THE TEAMS
SurrealDB gives us a foundation where we can unify semantic search, knowledge graphs, and AI-driven decision making without stitching together multiple systems. Collapsing responsibility into SurrealDB has become our default engineering posture.
VP of Engineering, Later