---
title: "Vector embeddings and search | SurrealDB University"
description: "What an embedding actually is, and how SurrealDB searches over one - exact search, the HNSW index, the KNN operator, and how to choose a model."
url: https://surrealdb.com/learn/ai/vector-embeddings
---

![Course content preview](https://surrealdb.com/assets/static/course-ai.LEsq_J_G.avif)

Course chapters

[Back to courses](https://surrealdb.com/learn) [SurrealDB for AI Engineers](https://surrealdb.com/learn/ai) [AI foundations](https://surrealdb.com/learn/ai/ai-foundations) [Vector embeddings and search](https://surrealdb.com/learn/ai/vector-embeddings) [Full-text search and BM25](https://surrealdb.com/learn/ai/fulltext-search-bm25) [Building a RAG knowledge base](https://surrealdb.com/learn/ai/rag-knowledge-base) [Hybrid search and reranking](https://surrealdb.com/learn/ai/hybrid-search-reranking) [Building an agent memory store](https://surrealdb.com/learn/ai/agent-memory-store) [Text-to-SurQL: agentic prompt engineering](https://surrealdb.com/learn/ai/text-to-surrealql) [Many sources and agents, one context layer](https://surrealdb.com/learn/ai/multi-source-context-layer) [Chunking strategies](https://surrealdb.com/learn/ai/chunking-strategies) [Evaluating retrieval quality](https://surrealdb.com/learn/ai/evaluating-retrieval) [Graph RAG beyond one hop](https://surrealdb.com/learn/ai/graph-rag-multi-hop) Certificate Pending

# Vector embeddings and search

**SurrealQL functions used here for the first time** 

- [`vector::distance::knn`](https://surrealdb.com/docs/reference/query-language/functions/database-functions/vector#vectordistanceknn) - the distance the `<|K, EF|>` operator already worked out, so the index is not asked twice
- [`vector::similarity::cosine`](https://surrealdb.com/docs/reference/query-language/functions/database-functions/vector#vectorsimilaritycosine) - scores two vectors against each other directly, with no index involved

## What an embedding is

An embedding is a **position**. A model reads something (a sentence, an image, a film synopsis) and returns a list of numbers giving that thing's coordinates in a space where *similar things sit close together*.

That's it, and it is easier to see in two dimensions than in dimensions like 1536 a real model might use. Here are six films, each one placed on two axes, **action** and **comedy**:

| Film | action | comedy |
| --- | --- | --- |
| John Wick | 0.98 | 0.08 |
| Die Hard | 0.92 | 0.20 |
| Shaun of the Dead | 0.45 | 0.90 |
| Airplane! | 0.10 | 0.98 |
| The Hangover | 0.10 | 0.95 |
| Manchester by the Sea | 0.02 | 0.05 |

```text
 comedy
   1.0 ┤ Airplane!
       │ The Hangover      Shaun of the Dead
       │
   0.5 ┤
       │
       │                                    Die Hard
   0.0 ┤ Manchester by the Sea                John Wick
       └──────────────────────────────────────────────
         0.0                                       1.0
                                                 action
```

Airplane! and The Hangover are almost on top of each other. Shaun of the Dead sits between the comedies and the action films, which is exactly right: it's both. Manchester by the Sea is off on its own, near nothing.

Now "find me something like *John Wick*" isn't about matching words any more. It's a geometry question: **which points are nearest to this point?** That has a number you can compute, which is why it works on things that share no words at all.

A real model changes three things, none of them the geometry:

- **The axes stop being readable.** A model gives you hundreds or thousands of dimensions, with 384, 768 and 1536 among the common sizes, and no single one of them means "action". Meaning is spread across all of them.
- **You stop choosing the numbers.** The model derives them from the content itself.
- **The space belongs to the model.** There is no canonical map of meaning that every model approximates a little better or worse. Each one carves up its own space during training, so the coordinates `text-embedding-3-small` gives a sentence bear no relation to the ones `bge-large` gives the same sentence. Two models can both be excellent and still place that sentence somewhere completely different. What survives across them is the *relative* geometry: whichever space you are in, John Wick should land nearer Die Hard than Airplane!

If you want more on this:

- [Embeddings: What they are and why they matter](https://simonwillison.net/2023/Oct/23/embeddings/): Simon Willison. Start here.
- [The Illustrated Word2vec](https://jalammar.github.io/illustrated-word2vec/): Jay Alammar. Where the "king − man + woman ≈ queen" intuition comes from.

## Step 1 - start a server

The `surreal` binary is all this lesson needs, and the [introduction](https://surrealdb.com/learn/ai) has the install command for both Bash and PowerShell. The vectors are hand-written 4-dimensional ones so that you can check the answers by eye, which means no embedding model and no cloud account.

```bash
surreal start --user root --pass secret
```

Leave that running in its own terminal. Started like this, the server runs on SurrealMX, the in-memory engine that has been the default since SurrealDB 3.0, so the database resets when you stop the process.

In-memory doesn't have to mean disposable, though, because SurrealMX can take periodic snapshots or run append-only persistence and still answer versioned queries. [Run a single-node, in-memory server](https://surrealdb.com/docs/running/in-memory) covers how to turn that on, and [File-backed storage](https://surrealdb.com/docs/running/file-backed) covers RocksDB when the data belongs on disk from the start. The rest of this track assumes the plain in-memory default, but feel free to choose a persistent option if you want to save your data.

## Step 2 - define the schema

This time we'll move from two to four axes (`[ action, comedy, sci-fi, drama ]`) because four is enough for films to disagree in interesting ways and still be readable.

```surql
DEFINE TABLE OVERWRITE film SCHEMAFULL;

DEFINE FIELD OVERWRITE title     ON film TYPE string;
DEFINE FIELD OVERWRITE year      ON film TYPE int;
DEFINE FIELD OVERWRITE embedding ON film TYPE array<float, 4>;

DEFINE INDEX OVERWRITE film_vec ON film
    FIELDS embedding
    HNSW DIMENSION 4
    DIST COSINE
    TYPE F32
    EFC 150 M 12 M0 24;
```

The `array<float, 4>` type is there for a reason. A vector of the wrong length is the usual way a vector table goes bad, and it usually happens quietly: a batch job embeds with a different model and the records just stop matching anything. A fixed-length array type rejects it at write time, and the field and the index's `DIMENSION 4` then state the same fact instead of agreeing by convention.

**HNSW** (*Hierarchical Navigable Small World*) is an approximate-nearest-neighbour index. It builds a layered graph of your vectors so that a search can hop towards the query instead of comparing against every record, which is what makes vector search viable over millions of records.

Only `DIMENSION` is required. Every other clause is a tuning knob with a default. They're written out in full above so you can see them once:

| Clause | Default | What it controls |
| --- | --- | --- |
| `DIMENSION` | *required* | The length of every embedding, which has to match your model's output exactly. |
| `DIST` | `EUCLIDEAN` | The distance function. Cosine compares *direction* and ignores magnitude, which is what text embeddings want. `MANHATTAN` is also available. |
| `TYPE` | `F32` | Storage precision. `F32` is the default because it matches what embedding models emit, and it halves the index size compared with `F64` for no accuracy that matters. The integer types are for quantised vectors. |
| `EFC` | `150` | Build-time search effort. Higher builds a better-connected graph, more slowly. The cost is paid once and the recall benefit is permanent. |
| `M` / `M0` | `12` / `24` | Links kept per node, and at the base layer. Higher improves recall and grows the index, with `M0` at twice `M` as the usual pairing. |

So of the five clauses above, four are just the defaults spelled out, while `DIST COSINE` is the only one that changes anything. These two definitions are identical as far as the engine is concerned:

```surql
DEFINE INDEX OVERWRITE film_vec ON film FIELDS embedding
    HNSW DIMENSION 4 DIST COSINE TYPE F32 EFC 150 M 12 M0 24;

DEFINE INDEX OVERWRITE film_vec ON film FIELDS embedding
    HNSW DIMENSION 4 DIST COSINE;
```

`INFO FOR TABLE film` shows the short form filled back out to the long one, along with an `LM` parameter that SurrealDB derives from the others and that is best left alone. So the short form is the one to use day to day, and the tuning starts when recall or build time turns into a number you can measure.

Before we finish this step, let's make sure we know what *quantising* mentioned above means. Quantising a vector means storing each component in a smaller type than the float the model produced, so the same 1536-dimensional embedding takes half the space at `I16` that it takes at `F32`. Less memory per vector means more of the index stays resident and each comparison is cheaper, paid for with a little recall, because two vectors that were only just distinguishable can round to the same value. It is a lever for corpora big enough that the index has stopped fitting comfortably, which is a long way past anything in this track.

Save the statements above as `schema.surql`, starting the file with `OPTION IMPORT;`, which `surreal import` requires. Then load it:

Bash

PowerShell

```bash
surreal import --endpoint http://localhost:8000 \
  --user root --pass secret --ns ai --db vectors schema.surql
```

## Step 3 - seed the films

25 films, hand-placed so you can check the distances yourself:

```surql
CREATE film:john_wick   SET title = "John Wick",         year = 2014, embedding = [0.98, 0.08, 0.02, 0.12];
CREATE film:matrix      SET title = "The Matrix",        year = 1999, embedding = [0.85, 0.08, 0.90, 0.20];
CREATE film:arrival     SET title = "Arrival",           year = 2016, embedding = [0.10, 0.02, 0.92, 0.85];
CREATE film:airplane    SET title = "Airplane!",         year = 1980, embedding = [0.10, 0.98, 0.05, 0.02];
CREATE film:godfather   SET title = "The Godfather",     year = 1972, embedding = [0.35, 0.05, 0.02, 0.95];
-- ...and twenty more, in full below
```

**The complete `seed.surql` - all 25 films** 

```surql
OPTION IMPORT;

-- Action
CREATE film:mad_max     SET title = "Mad Max: Fury Road",     year = 2015, embedding = [0.97, 0.05, 0.25, 0.15];
CREATE film:john_wick   SET title = "John Wick",              year = 2014, embedding = [0.98, 0.08, 0.02, 0.12];
CREATE film:die_hard    SET title = "Die Hard",               year = 1988, embedding = [0.92, 0.20, 0.02, 0.15];
CREATE film:the_raid    SET title = "The Raid",               year = 2011, embedding = [0.99, 0.02, 0.02, 0.10];

-- Action + sci-fi
CREATE film:matrix      SET title = "The Matrix",             year = 1999, embedding = [0.85, 0.08, 0.90, 0.20];
CREATE film:terminator2 SET title = "Terminator 2",           year = 1991, embedding = [0.88, 0.10, 0.85, 0.25];
CREATE film:edge        SET title = "Edge of Tomorrow",       year = 2014, embedding = [0.85, 0.30, 0.88, 0.15];
CREATE film:aliens      SET title = "Aliens",                 year = 1986, embedding = [0.85, 0.05, 0.85, 0.20];

-- Sci-fi + drama
CREATE film:br2049      SET title = "Blade Runner 2049",      year = 2017, embedding = [0.30, 0.02, 0.95, 0.75];
CREATE film:arrival     SET title = "Arrival",                year = 2016, embedding = [0.10, 0.02, 0.92, 0.85];
CREATE film:interstellar SET title = "Interstellar",          year = 2014, embedding = [0.25, 0.05, 0.93, 0.80];
CREATE film:solaris     SET title = "Solaris",                year = 1972, embedding = [0.02, 0.02, 0.88, 0.90];

-- Comedy
CREATE film:airplane    SET title = "Airplane!",              year = 1980, embedding = [0.10, 0.98, 0.05, 0.02];
CREATE film:hangover    SET title = "The Hangover",           year = 2009, embedding = [0.10, 0.95, 0.02, 0.10];
CREATE film:groundhog   SET title = "Groundhog Day",          year = 1993, embedding = [0.05, 0.88, 0.30, 0.35];
CREATE film:shaun       SET title = "Shaun of the Dead",      year = 2004, embedding = [0.45, 0.90, 0.25, 0.10];

-- Comedy + sci-fi
CREATE film:bttf        SET title = "Back to the Future",     year = 1985, embedding = [0.35, 0.80, 0.85, 0.15];
CREATE film:ghostbusters SET title = "Ghostbusters",          year = 1984, embedding = [0.40, 0.88, 0.70, 0.05];
CREATE film:galaxy_quest SET title = "Galaxy Quest",          year = 1999, embedding = [0.35, 0.85, 0.80, 0.10];

-- Drama
CREATE film:godfather   SET title = "The Godfather",          year = 1972, embedding = [0.35, 0.05, 0.02, 0.95];
CREATE film:manchester  SET title = "Manchester by the Sea",  year = 2016, embedding = [0.02, 0.05, 0.02, 0.98];
CREATE film:twelve_men  SET title = "12 Angry Men",           year = 1957, embedding = [0.05, 0.05, 0.02, 0.96];
CREATE film:whiplash    SET title = "Whiplash",               year = 2014, embedding = [0.15, 0.05, 0.02, 0.95];

-- Action + drama
CREATE film:heat        SET title = "Heat",                   year = 1995, embedding = [0.80, 0.05, 0.02, 0.75];
CREATE film:gladiator   SET title = "Gladiator",              year = 2000, embedding = [0.85, 0.05, 0.05, 0.70];
```

Bash

PowerShell

```bash
surreal import --endpoint http://localhost:8000 \
  --user root --pass secret --ns ai --db vectors seed.surql
```

## Step 4 - search with no index at all

This first search uses no index at all, which is what the index will speed up. Open the shell and pipe in the query file:

Bash

PowerShell

```bash
surreal sql --endpoint ws://localhost:8000 \
  --user root --pass secret --ns ai --db vectors --pretty < queries.surql
```

Our query vector points at the action/sci-fi corner, roughly *"an action sci-fi film"*:

```surql
LET $q = [0.85, 0.10, 0.90, 0.20];
```

**The complete `queries.surql` - the five demos below, one statement per line** 

```surql
-- The query vector: roughly "an action sci-fi film", the neighbourhood The Matrix sits in.
LET $q = [0.85, 0.10, 0.90, 0.20];

-- 1. Exact search, no index involved: compare against every record and sort.
SELECT title, year, vector::similarity::cosine(embedding, $q) AS similarity FROM film ORDER BY similarity DESC LIMIT 5;

-- 2. The same answer through the HNSW index. Note the direction flip: knn() returns distance, so smaller is closer.
SELECT title, year, vector::distance::knn() AS dist FROM film WHERE embedding <|5, 40|> $q ORDER BY dist;

-- 3. Distance back to a 0-1 similarity score. With a COSINE index, score = 1 - distance.
SELECT title, score FROM (SELECT title, (1 - vector::distance::knn()) AS score FROM film WHERE embedding <|5, 40|> $q) ORDER BY score DESC;

-- 4. EF below K silently returns fewer than K records - asked for 5, explored 2.
SELECT title, vector::distance::knn() AS dist FROM film WHERE embedding <|5, 2|> $q ORDER BY dist;

-- 5. Ten nearest, which on this dataset matches an exact full scan exactly.
SELECT title, vector::distance::knn() AS dist FROM film WHERE embedding <|10, 40|> $q ORDER BY dist;
```

> **On running the files.** `queries.surql` keeps each statement on a single line, because a file piped into `surreal sql` doesn't treat line breaks inside a statement as whitespace. The queries are formatted below for readability, and in the interactive shell a wrapped line ends with a `\`. Loading through `surreal import` has no such restriction.

`vector::similarity::cosine()` compares two vectors directly, with no index involved: it reads every record, scores it, and sorts:

```surql
SELECT title, year, vector::similarity::cosine(embedding, $q) AS similarity
FROM film
ORDER BY similarity DESC
LIMIT 5;
```

Output

```surql
[
    { title: 'The Matrix',         year: 1999, similarity: 0.99987 },
    { title: 'Aliens',             year: 1986, similarity: 0.99885 },
    { title: 'Terminator 2',       year: 1991, similarity: 0.99814 },
    { title: 'Edge of Tomorrow',   year: 2014, similarity: 0.98659 },
    { title: 'Mad Max: Fury Road', year: 2015, similarity: 0.85011 }
]
```

Cosine similarity runs from `1.0` (same direction) to `-1.0` (opposite). The four action-sci-fi films score above `0.98`; Mad Max, which is action but barely sci-fi, drops to `0.85`. Nothing else comes close.

That's vector search, and on 25 records it's instant. You don't need an index to start, you need one when a full scan stops being free.

## Step 5 - the same search, through the index

The `<|K, EF|>` operator hands the query to HNSW instead. `K` is how many neighbours to return, `EF` is the search effort:

```surql
SELECT title, year, vector::distance::knn() AS dist
FROM film
WHERE embedding <|5, 40|> $q
ORDER BY dist;
```

Output

```surql
[
    { title: 'The Matrix',         year: 1999, dist: 0.00013 },
    { title: 'Aliens',             year: 1986, dist: 0.00115 },
    { title: 'Terminator 2',       year: 1991, dist: 0.00186 },
    { title: 'Edge of Tomorrow',   year: 2014, dist: 0.01341 },
    { title: 'Mad Max: Fury Road', year: 2015, dist: 0.14989 }
]
```

Same five films, same order. Two differences in how you read it:

- **`vector::distance::knn()` returns distance, not similarity**, so **smaller is closer**, the opposite direction to Step 4. It reports whatever distance function the index was defined with, which is why you don't name cosine anywhere in this query.
- **You don't pass the vector to a function.** The operator does the comparison; `knn()` just reads back the number the index already computed. This only works on a field with a matching vector index.

Since the index is cosine, `1 - distance` converts straight back to the similarity from Step 4:

```surql
SELECT title, score FROM (
    SELECT title, (1 - vector::distance::knn()) AS score
    FROM film
    WHERE embedding <|5, 40|> $q
)
ORDER BY score DESC;
```

Output

```surql
[
    { title: 'The Matrix',         score: 0.99987 },
    { title: 'Aliens',             score: 0.99885 },
    { title: 'Terminator 2',       score: 0.99814 },
    { title: 'Edge of Tomorrow',   score: 0.98659 },
    { title: 'Mad Max: Fury Road', score: 0.85011 }
]
```

Identical to the exact scores in Step 4, to five decimal places. A 0-1 score is much easier to reason about than a distance, and it's what you set thresholds on. [Lesson 04](https://surrealdb.com/learn/ai/rag-knowledge-base) uses this to keep weak matches out of a prompt.

## Step 6 - the "approximate" in approximate nearest neighbour

HNSW is not guaranteed to return the true nearest neighbours. It explores part of the graph and returns the best it found, and `EF` is the size of that exploration. Set it too low and the search gives up early:

```surql
-- Asking for 5, but only exploring 2
SELECT title, vector::distance::knn() AS dist
FROM film
WHERE embedding <|5, 2|> $q
ORDER BY dist;
```

Output

```surql
[
    { title: 'The Matrix', dist: 0.00013 },
    { title: 'Aliens',     dist: 0.00115 }
]
```

This gave us two results, not five, and no errors or warnings. The query just returned less than it was asked for, because `EF` caps how many candidates the search will consider. On this dataset the results you *do* get are correct; you just silently get fewer of them. In an application that looks like a thin prompt, not a broken query: the sort of bug that ships.

**Keep `EF` comfortably above `K`.** `40` is a sensible default, and is what the rest of this course uses.

With adequate effort, the index agrees with the full scan completely:

```surql
SELECT title, vector::distance::knn() AS dist
FROM film
WHERE embedding <|10, 40|> $q
ORDER BY dist;
```

Output

```surql
[
    { title: 'The Matrix',         dist: 0.00013 },
    { title: 'Aliens',             dist: 0.00115 },
    { title: 'Terminator 2',       dist: 0.00186 },
    { title: 'Edge of Tomorrow',   dist: 0.01341 },
    { title: 'Mad Max: Fury Road', dist: 0.14989 },
    { title: 'Blade Runner 2049',  dist: 0.19561 },
    { title: 'Interstellar',       dist: 0.22947 },
    { title: 'Back to the Future', dist: 0.24088 },
    { title: 'Galaxy Quest',       dist: 0.27022 },
    { title: 'Ghostbusters',       dist: 0.28927 }
]
```

Those are the same ten films, in the same order, as a full exact scan of the table.

That match says less than it appears to. On 25 vectors in 4 dimensions HNSW visits nearly everything, so agreeing with the exact scan is what the size of the data guarantees rather than evidence the index is accurate. The approximation only starts costing you recall at scale: millions of records, hundreds of dimensions, a graph the search can't walk in full. That's when tuning `EFC`, `M` and `EF` starts to matter, and when measuring recall against a sample whose correct answers you already know does too.

The tail of the list shows the same thing. Blade Runner 2049 and Interstellar are sci-fi *dramas*; the query asked for action sci-fi, so they rank behind everything action-flavoured but well ahead of Airplane! or The Godfather, which don't appear at all. The ranking falls off smoothly, which is why semantic search still helps when nothing is an exact match.

## Choosing an embedding model

The lesson so far used numbers picked by hand. In production a model picks them, and that choice has consequences reaching into your schema.

That there is no shared space is also why an embedding model market exists at all. If one true set of coordinates were out there to be found, embeddings would be a commodity and the only contest would be who computed them fastest. Instead each model trains its own space, and they compete on how cleanly that space separates the things a particular job cares about. That's why leaderboards break down by task and by language rather than reporting a single number, and why a model that leads on code retrieval can sit mid-table on multilingual search. Some vendors keep the weights closed and sell access to them; others publish the weights and compete on quality. [Our comparison of OpenAI, Google, Qwen, Nomic, Jina and BAAI](https://surrealdb.com/blog/embedding-models-comparison) works through what that means for a specific shortlist.

Some notes when choosing an embedding model for SurrealDB:

1. \1.

   **`DIMENSION` must equal the model's output.** `text-embedding-3-small` gives 1536, `all-MiniLM-L6-v2` gives 384, `bge-large` gives 1024. Set it once, and `array<float, 1536>` on the field catches a mismatch on write, as long as the two are changed together.
2. \2.

   **The metric follows the model.** Most text embedding models are trained for cosine similarity, which is why `DIST COSINE` is the default choice here. Your model's own documentation settles that, before anyone chooses `EUCLIDEAN`.
3. \3.

   **Bigger vectors cost on every axis.** 1536 dimensions is four times the index size and four times the per-comparison work of 384. `TYPE F32` halves it back versus `F64` for no meaningful accuracy loss. Some models also support *Matryoshka* truncation (using the first N dimensions of a longer vector) which trades a little accuracy for a lot of space.
4. \4.

   **Changing models means re-embedding everything.** Vectors from two different models are not comparable, even at the same dimensionality, the axes mean different things. A model change is a re-embed of the whole corpus plus an index rebuild, so benchmark on your own data before you commit.
5. \5.

   **The same model has to be used at write time and at query time.** Embedding your documents with one model and your queries with another produces confidently wrong results rather than errors.

The [MTEB leaderboard](https://huggingface.co/spaces/mteb/leaderboard) ranks models across retrieval benchmarks and is the usual starting point, but the benchmark that matters is yours.

## Up next in the course

We've now covered the building block for the rest of the course: text in, vector out, nearest neighbours back. It's very good at meaning and, as you'll see, quite bad at some things that look easy.

Ask this index about an error code, an order number, or a product name and it will struggle, because a bare token like `429` carries almost no semantic signal, so its vector points nowhere useful. [Lesson 03](https://surrealdb.com/learn/ai/fulltext-search-bm25) covers that with full-text search and BM25, using these same 25 films, so you can watch the two approaches succeed and fail on identical data. [Lesson 04](https://surrealdb.com/learn/ai/rag-knowledge-base) builds a knowledge base on the semantic search you have just seen, and [lesson 05](https://surrealdb.com/learn/ai/hybrid-search-reranking) runs both retrievers side by side and fuses their rankings.

At a glance

**You are on**

Chapter 2 of 11

**Chapters**

11

**Format**

Text, with queries you can run

**Runs in**

SurrealDB Studio, in the browser

**Cost**

Free

**Certificate**

On completion

[Next: Full-text search and BM25](https://surrealdb.com/learn/ai/fulltext-search-bm25)

## Continue

### [AI foundations](https://surrealdb.com/learn/ai/ai-foundations)

Previous

### [Full-text search and BM25](https://surrealdb.com/learn/ai/fulltext-search-bm25)

Next lesson

THE PLATFORM

## Everything an application and its agents know. Five surfaces, one engine.

Database

Document, graph, vector, time-series and relational in one engine.

![Five data models as dotted tiles: documents, graph, vector, time-series and relational](https://surrealdb.com/assets/static/platform-database.DUdumYDz.avif)

Read more

[Database](https://surrealdb.com/surrealdb)

Agent Memory

What an agent learns, with its source and its time, in the same engine.

![A timeline of remembered facts, each with its source](https://surrealdb.com/assets/static/platform-agent-memory.B4RNjvbX.avif)

Read more

[Agent Memory](https://surrealdb.com/agent-memory)

Cloud

Managed clusters in the regions you choose, scaled on demand.

![Clusters in three regions on a world map, each running or scaling](https://surrealdb.com/assets/static/platform-cloud.--QZnaVi.avif)

Read more

[Cloud](https://surrealdb.com/cloud)

Studio

Query, explore and design the schema from the browser.

![A SurrealQL query in Studio and the schema graph under it](https://surrealdb.com/assets/static/platform-studio.7ykBFNLq.avif)

Read more

[Studio](https://surrealdb.com/studio)

MCP

Every model that speaks MCP reaches the database and the memory directly.

![Three models connected through MCP to the database and Agent Memory](https://surrealdb.com/assets/static/platform-mcp.D_oH0_wm.avif)

Read more

[MCP](https://surrealdb.com/mcp)

IN PRODUCTION

## Trusted at scale. Samsung, Nvidia, Verizon, Tencent, and Walmart run on SurrealDB.

14,000+

Developers building on SurrealDB Cloud

4M+

Developers building on SurrealDB worldwide

FROM THE TEAMS

> SurrealDB gives us a foundation where we can unify semantic search, knowledge graphs, and AI-driven decision making without stitching together multiple systems. Collapsing responsibility into SurrealDB has become our default engineering posture.

*Justin Foley*

VP of Engineering, Later

```json
{"@context":"https://schema.org","@type":"Course","name":"SurrealDB for AI Engineers","description":"An eleven-lesson course from calling an LLM API to an agent that retrieves by meaning, remembers what it learned, writes its own queries, and can prove its retrieval works - all against one database.","url":"https://surrealdb.com/learn/ai","inLanguage":"en","isAccessibleForFree":false,"provider":{"@type":"Organization","name":"SurrealDB","url":"https://surrealdb.com"},"hasPart":[{"@type":"LearningResource","name":"SurrealDB for AI Engineers","url":"https://surrealdb.com/learn/ai"},{"@type":"LearningResource","name":"AI foundations","url":"https://surrealdb.com/learn/ai/ai-foundations"},{"@type":"LearningResource","name":"Vector embeddings and search","url":"https://surrealdb.com/learn/ai/vector-embeddings"},{"@type":"LearningResource","name":"Full-text search and BM25","url":"https://surrealdb.com/learn/ai/fulltext-search-bm25"},{"@type":"LearningResource","name":"Building a RAG knowledge base","url":"https://surrealdb.com/learn/ai/rag-knowledge-base"},{"@type":"LearningResource","name":"Hybrid search and reranking","url":"https://surrealdb.com/learn/ai/hybrid-search-reranking"},{"@type":"LearningResource","name":"Building an agent memory store","url":"https://surrealdb.com/learn/ai/agent-memory-store"},{"@type":"LearningResource","name":"Text-to-SurQL: agentic prompt engineering","url":"https://surrealdb.com/learn/ai/text-to-surrealql"},{"@type":"LearningResource","name":"Many sources and agents, one context layer","url":"https://surrealdb.com/learn/ai/multi-source-context-layer"},{"@type":"LearningResource","name":"Chunking strategies","url":"https://surrealdb.com/learn/ai/chunking-strategies"},{"@type":"LearningResource","name":"Evaluating retrieval quality","url":"https://surrealdb.com/learn/ai/evaluating-retrieval"},{"@type":"LearningResource","name":"Graph RAG beyond one hop","url":"https://surrealdb.com/learn/ai/graph-rag-multi-hop"}]}
```

```json
{"@context":"https://schema.org","@type":"LearningResource","name":"Vector embeddings and search","description":"What an embedding actually is, and how SurrealDB searches over one - exact search, the HNSW index, the KNN operator, and how to choose a model.","url":"https://surrealdb.com/learn/ai/vector-embeddings","learningResourceType":"lesson","isPartOf":{"@type":"Course","name":"SurrealDB for AI Engineers","url":"https://surrealdb.com/learn/ai"},"position":3}
```

```json
{"@context":"https://schema.org","@type":"Organization","@id":"https://surrealdb.com/#organization","name":"SurrealDB","url":"https://surrealdb.com","logo":"https://surrealdb.com/assets/static/logo.BG7_TG2b.svg","description":"SurrealDB is the context and memory layer for AI agents. A multi-model database for documents, graphs, vectors, and time-series.","foundingDate":"2022","legalName":"SurrealDB Ltd","identifier":{"@type":"PropertyValue","propertyID":"GB-COH","value":"13615201"},"address":{"@type":"PostalAddress","streetAddress":"3rd Floor, 1 Ashley Road","addressLocality":"Altrincham","addressRegion":"Cheshire","postalCode":"WA14 2DT","addressCountry":"GB"},"contactPoint":[{"@type":"ContactPoint","contactType":"customer support","email":"support@surrealdb.com","url":"https://surrealdb.com/contact","availableLanguage":"English"},{"@type":"ContactPoint","contactType":"sales","email":"info@surrealdb.com","url":"https://surrealdb.com/contact","availableLanguage":"English"},{"@type":"ContactPoint","contactType":"security","email":"security@surrealdb.com","url":"https://surrealdb.com/.well-known/security.txt","availableLanguage":"English"},{"@type":"ContactPoint","contactType":"legal","email":"legal@surrealdb.com","url":"https://surrealdb.com/legal","availableLanguage":"English"}],"hasCertification":[{"@type":"Certification","name":"SOC 2 Type 2"},{"@type":"Certification","name":"GDPR"},{"@type":"Certification","name":"Cyber Essentials Plus"},{"@type":"Certification","name":"ISO 27001"}],"owns":[{"@type":"SoftwareApplication","name":"SurrealDB","url":"https://surrealdb.com/surrealdb"},{"@type":"SoftwareApplication","name":"Agent Memory","url":"https://surrealdb.com/agent-memory"}],"knowsAbout":["multi-model databases","document databases","graph databases","vector search","time-series databases","SurrealQL","Agent Memory","real-time databases","embedded databases","context layer","graph ontology","distributed database","knowledge graphs","distributed transaction protocols","highly-scalable databases"],"sameAs":["https://www.wikidata.org/wiki/Q124316308","https://github.com/surrealdb/surrealdb","https://twitter.com/surrealdb","https://www.youtube.com/@surrealdb","https://www.linkedin.com/company/surrealdb","https://discord.gg/surrealdb","https://www.reddit.com/r/surrealdb","https://www.instagram.com/surrealdb","https://medium.com/surrealdb","https://dev.to/surrealdb"]}
```

```json
{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://surrealdb.com"},{"@type":"ListItem","position":2,"name":"Learn","item":"https://surrealdb.com/learn"},{"@type":"ListItem","position":3,"name":"Vector embeddings","item":"https://surrealdb.com/learn/ai/vector-embeddings"}]}
```
