Skip to content
Course content preview

AI foundations

This first lesson is the background the rest of the course uses: how a language model works, and the difference between writing a prompt and choosing what goes into the window. The reading is all external. Lessons 02 to 08 assume you've seen these ideas.

The later lessons use the usual terms - context window, context rot, attention budget, the agent loop - and put a database underneath them. Each link below is here because a later lesson needs it, and the note says which one.

A language model does one thing: given a sequence of tokens, predict the next one.

Everything else (answering questions, writing code, holding a conversation) is that operation in a loop, each prediction fed back in.

Two things follow from that, and both show up for the rest of the course:

  1. 1.

    The context window is the model's entire world at inference time: no memory between calls, no access to your data, so a fact that isn't in the window doesn't exist.

  2. 2.

    Nothing in predict a plausible next token checks whether the continuation is true - hallucination is what the mechanism does when the window doesn't hold the answer. The fix is to put the answer there.

Resource

What it covers

Transformers, the tech behind LLMs

3Blue1Brown, 27 minutes

Grant Sanderson builds a transformer up one piece at a time, animating what an embedding, an attention head and a softmax are each doing to the numbers. If you watch only one thing here, watch this one: lesson 02's talk of vectors and closeness in meaning gets a lot less abstract once you have seen the geometry move.

Intro to Large Language Models

Andrej Karpathy, one hour

Pretraining, fine-tuning and where capabilities actually come from, at a brisk pace. Karpathy is blunt about what models can't do, which most intros skip. He also draws a line this course sticks to: capability comes from training, while knowledge of your data has to arrive at inference time. That's why nothing from lesson 02 onwards fine-tunes a model.

The Illustrated Transformer

Jay Alammar

The diagram set that everyone else borrows from. It covers similar ground to the 3Blue1Brown video but on the page, which is easier when you want to stop on one step and stare at it. Its walk through how a token becomes a vector is the piece lesson 02 builds on most directly.

Tiktokenizer

Tat Dat Duong

A tool rather than a reading. Paste in any text and watch it break into tokens. Two minutes here makes the context window concrete. Notice how little a bare error code or an order number gives a model to work with: those are the queries lesson 03 is for.

The Tiktokenizer web app tokenising four lines of text, each twenty characters long, with the coloured token blocks growing more numerous down the four lines and a total token count of 29

Four lines, twenty characters each, run through gpt-4o. The token count climbs with unfamiliarity rather than with length: I like cats and dogs costs 5, the same sentence with SurrealDB in it costs 6, an error code and an order number cost 7, and an invented word costs 8. Every step down that list leaves a model less to work with, which is the blind spot lesson 03 picks up. Screenshot of Tiktokenizer, built by Tat Dat Duong and MIT licensed.

Prompt engineering is the wording of the instruction: role, format, examples. Write it once and reuse it.

Context engineering is deciding what information occupies the finite window on each request, and where it comes from. You solve that at runtime, on every call, against data that changes. It's a retrieval problem, which makes it a database problem.

Lessons 04 to 07 are all context engineering: retrieving the right documents, fusing two rankings, recalling what the agent learned last week, and generating the prompt itself out of live schema so it can't go stale.

When the data keeps changing. A store that only ever grows ends up holding the old version of a fact next to the new one, and a retriever with no idea of time will hand a model both. SurrealDB's Agent Memory is built around that: it keeps what was said, what is true now, and what used to be true as three separate things, so a superseded fact can be closed rather than left to compete with its replacement. Lesson 06 builds a small version of the same idea by hand, with an expiry field and a write-back on recall.

Resource

What it covers

Effective context engineering for AI agents

Anthropic, September 2025

The distinction set out in full by the people who popularised the term. The context rot section is the useful bit: past a certain size, adding tokens to the window measurably degrades the model's recall, so you want less in the window, not more. It also covers compaction, structured note-taking and sub-agent architectures as three ways of living within an attention budget. Lesson 04 turns that into a WHERE clause: weak matches never reach the prompt.

Prompt engineering overview

Anthropic

The practical checklist: be explicit, give examples, let the model think, assign it a role. Skim it once now and come back to it when a prompt misbehaves. It works better as a reference than as a read. Lesson 07 uses this checklist almost line by line, except the database assembles the prompt rather than a person.

Building effective AI agents

Anthropic, December 2024

Where retrieval sits inside an agent loop, and the argument for simple composable patterns over frameworks. The post now opens with a warning that its tooling advice has dated; the patterns have held up rather better than the tools around them. The agent loop it describes is the one lesson 06 writes back into, and the multi-agent shape it sketches is what lesson 08 takes apart.

LLM powered autonomous agents

Lilian Weng, about half an hour's read

A survey, organised as planning, memory and tool use. The memory section is the bit lesson 06 follows most closely, and its detour through maximum inner product search is the same retrieval problem that lesson 02 hands to an HNSW index.

The window is finite, and nobody gets to choose what the user asks. So the question is: given a question you never anticipated, what goes in the window?

There are two ways to find that content, by meaning and by literal term. Lesson 02 covers the first, lesson 03 the second, and every later lesson leans on one or both.

THE PLATFORM

Everything an application and its agents know. Five surfaces, one engine.

IN PRODUCTION

Trusted at scale. Samsung, Nvidia, Verizon, Tencent, and Walmart run on SurrealDB.

14,000+

Developers building on SurrealDB Cloud

4M+

Developers building on SurrealDB worldwide

FROM THE TEAMS

SurrealDB gives us a foundation where we can unify semantic search, knowledge graphs, and AI-driven decision making without stitching together multiple systems. Collapsing responsibility into SurrealDB has become our default engineering posture.
Justin Foley

VP of Engineering, Later

SurrealDB

The context and memory layer for AI agents

Database. Graphs, vectors, documents and relational data in one engine, in a single ACID transaction.
Agent Memory. Connects and retrieves context wherever your data lives, every fact carrying its source.
Cloud. Fully managed, in the cloud provider and region you choose.

Explore with AI

Copyright © 2026 SurrealDB Ltd. Registered in England and Wales. Company no. 13615201

Registered address: 3rd Floor 1 Ashley Road, Altrincham, Cheshire, WA14 2DT, United Kingdom

Trading address: Huckletree Oxford Circus, 213 Oxford Street, London, W1D 2LG, United Kingdom