> Full SurrealDB documentation index: https://surrealdb.com/docs/llms.txt

# Vulcan, Alberta: results

Scores and costs from the Vulcan, Alberta demo: Claude Haiku 4.5, Sonnet 5 and Opus 5.5 answering 19 questions about a Wikipedia article's history with today's page, the page and its history, or SurrealDB Agent Memory.

This page gives the scores and costs from the [Vulcan, Alberta demo](/docs/agent-memory/cookbooks/showcase/vulcan.md). Three Claude models, each with three sets of tools, answered the same 19 questions about one Wikipedia article and its history, three times each. This is a single demo on one article, meant to give a general idea of what Agent Memory adds, not a benchmark.

## A team's path to Agent Memory

The agents in this demo can be read as the steps a team might take when attempting to improve accuracy of its agents and eventually making the decision to move to SurrealDB Agent Memory. The steps below give the score out of 38 and the cost per question:

1. **Start with a standard model.** Sonnet 5 reading today's page scores 15.0, at $0.049 a question. Most of the questions are about what the page used to say, so the answers are not on it.
2. **Try a larger model.** Opus 5.5 with today's page costs a third more, $0.065 a question, and scores 14.0. A larger model often gives better answers, but it cannot answer from text that it does not have.
3. **Give the agent the page history.** Sonnet 5 reading the page and its history scores 25.7, at $0.096 a question, almost twice the first cost. The history is a list of raw revisions, so the agent has to decide which revisions to open, and it has to work out for itself which version of a claim is current and when it changed.
4. **Try Agent Memory with a smaller model.** Haiku 4.5 with Agent Memory scores 29.7, at $0.061 a question. That is higher than Sonnet 5 with the history, for about two thirds of the cost. Agent Memory returns the claims from every revision about a topic, with the dates that the article stated them, so the smaller model does not have to search for them.

The embedded page below compares all nine agents:

- **The chart** places each agent by its score and its cost per question, with one line for each set of tools.
- **The table under it** gives every question, with each agent's mean score over three runs. Select a question to read its answers.
- **The answers** show what each agent said, the tools it called and its score. Use the model buttons to switch between models, and the run buttons to see how the answers vary between runs.

<DemoEmbed title="Vulcan, Alberta: results" height={760} pages={[{ "label": "Three agents, one question", "src": "/docs/showcase/vulcan/agents.html" }]} />

## Accuracy and cost

Each answer is scored from 0 to 2 against a written answer key, by a model that does not know which agent wrote it, so 38 is a perfect score over the 19 questions. Cost is the API cost of an answer, including every tool call the agent makes. The agents that read Wikipedia fetch pages with the fetch tool of Claude Code, which hands the agent a summary of each page made by a smaller model, Haiku 4.5, rather than the whole page. That keeps their token counts and costs lower, and may be part of the reason why all three models score about the same with today's page.

The agents with today's page are a baseline, showing which questions need the history at all. They do not have the information to answer the rest, so the fair comparison for Agent Memory is the page and its history, which holds everything that Agent Memory holds.

| Tools | Model | Score out of 38, mean of 3 runs | Cost per question |
| --- | --- | --- | --- |
| Agent Memory | Haiku 4.5 | 29.7 | $0.061 |
| Agent Memory | Sonnet 5 | 32.0 | $0.179 |
| Agent Memory | Opus 5.5 | 35.7 | $0.357 |
| Page + history | Haiku 4.5 | 21.7 | $0.064 |
| Page + history | Sonnet 5 | 25.7 | $0.096 |
| Page + history | Opus 5.5 | 28.7 | $0.139 |
| Today's page | Haiku 4.5 | 15.0 | $0.039 |
| Today's page | Sonnet 5 | 15.0 | $0.049 |
| Today's page | Opus 5.5 | 14.0 | $0.065 |

- **Agent Memory raises the score at every model size.** It adds 8.0 points over the history with Haiku 4.5, 6.3 with Sonnet 5 and 7.0 with Opus 5.5.
- **With a small model, Agent Memory costs no more than the history.** Haiku 4.5 costs $0.061 a question with Agent Memory and $0.064 reading the history, and scores 8 points more. With the larger models, Agent Memory costs more per question. An agent with Agent Memory reads about three times as many tokens as one reading the history, and a larger model charges more for each token.
- **A small model with Agent Memory scores at least as well as a large model without it.** Haiku 4.5 with Agent Memory scores 29.7, against 28.7 for Opus 5.5 with the history, at $0.061 against $0.139 a question.
- **A larger model with Agent Memory adds more accuracy, at a higher cost.** Opus 5.5 with Agent Memory scores 35.7, the highest of all, for $0.357 a question. Use it where those points are worth the cost. Elsewhere, Haiku 4.5 with Agent Memory gives about 83% of that score for about a sixth of the price.
- **A gap of a point or two is a tie.** Scores vary between runs: Haiku 4.5 with Agent Memory scored 28, 32 and 29 in its three runs, and the table gives the mean.

## The clearest example

Agent Memory helps on most questions about what the article used to say. The clearest comparison is Haiku 4.5 with Agent Memory against Sonnet 5 reading the page and its history: the smaller model with Agent Memory costs less per question, and it scores higher on seven of the 19 questions, all of them about what the article used to say. [The answers side by side](/docs/agent-memory/cookbooks/showcase/vulcan/answers.md) shows three of them in full, with the part of the Agent Memory result that held each answer.

![A dark street of low buildings and telegraph poles, with a tornado funnel reaching down from a pale, swirling sky at the end of it.](~/assets/img/spectron/vulcan/tornado-1927.webp)

*The tornado approaching Vulcan on the evening of 8 July 1927. Until January 2010 the article gave the year as 1926. Since then it has said that this photograph was used for the "tornado" article in Encyclopaedia Britannica, a claim marked as needing a citation from 2020 until a source was added in 2022. Photo by McDermid Photo Laboratories, via [Wikimedia Commons](https://commons.wikimedia.org/wiki/File:Vulcan,_Alberta_1927.jpg), public domain.*

## The questions

The 19 questions fall into six groups, and each has a written answer key with the revision that supports it.

| Group | What it asks | Example |
| --- | --- | --- |
| Current facts | What today's page says. A control: every agent should score well. | What is Vulcan's population? |
| Change over time | How a value changed between revisions. | When did the tornado hit Vulcan? |
| Content that was removed | Claims that are no longer on the page. | How many grain elevators did Vulcan have, and are any left? |
| Stability and controversy | Which sections were edited most, and what was reverted. | Which claims in the History section were questioned, and what happened to them? |
| Talk page | Discussion on the article's talk page. | What have editors discussed on the talk page? |
| Negative control | A fact that the article has never held. The answer should say so. | What is Vulcan's sister city? |

An agent with today's page can answer the current facts and little else. An agent with the history can open old revisions, but it has to choose which of the 428 to open. The agent with Agent Memory asks for a topic, such as "grain elevators Vulcan Alberta", and receives the claims from every revision that mentioned it, with their dates.
