Skip to content

Showcase

/

Robin Hood character graph

How the Robin Hood graph was built

This page describes the pipeline that produced the Robin Hood character graph, in the order the steps run. Each step is a short script, and the same steps work for any book or other text that arrives in parts. Build your own character graph gives the API calls to reproduce it.

The text is Project Gutenberg's #10148: 22 chapters and about 107,000 words, and public domain, so the pages can ship with it. A full ingest takes about three hours, which is short enough to repeat after a change and compare the results.

A script reads the Gutenberg HTML and writes one clean text file per chapter, split on the chapter headings.

Each chapter is uploaded as a document and stamps it with observedAt, the time the memory treats as when the facts were learned. Chapter 1 is dated 1 January 2000, chapter 2 the next day, and so on, so the book reads as one day per chapter, and a reading position in the demo is exactly that date passed as asOf.

The chapters go in one at a time, in order. Extraction is shown the entities the Context already holds and told to reuse their names, so the order decides which form of a name later mentions attach to. After each chapter reports ready, the upload waits a little longer, because reconciliation lands shortly afterwards, and a chapter that starts before its predecessor's entities are committed mints fresh name variants instead of reusing them.

Important

Before a long ingest, upload one chapter and confirm that GET /entities is not empty. If extraction is not running, every later chapter still reports ready with nothing extracted, and the problem only shows at the end.

The next step reads every fact in the Context and groups them by the time they became known, which observedAt set to narrative time. That gives, for each entity, what was known about it at each reading position, and a dot's size on the graph is the number of facts known so far. The data comes from each entity's history, GET /entities/{type}/{name}/history, which returns every value with the time it became known and, for a replaced value, the time it stopped being current.

A script asks Gemini 2.5 Flash for a short description of each entity for each chapter, from the harvested facts and nothing else. Agent Memory plays no part in this step beyond supplying the facts. A description can only say what the memory captured, which is what makes it a fair picture of the memory rather than of the book.

A script suggests which names belong to one person, using two rules on entities of the same type:

RuleExample
The names differ only by an articlehost and the host
Every capitalised word of the shorter name is in the longer oneTuck and Friar Tuck, merry Robin and Robin Hood

A name that fits two unrelated longer names is left unlinked. Each group takes the name with the most capitalised words as its main name, and the suggestions go into the page's data with the rule that produced each one.

PageContent comes from
GraphThe harvested graph, descriptions and alias suggestions
IntroductionWritten text, with every number filled in from the book and the graph
What the memory returnsThe demo's own API calls, recorded with their responses and timings

Each page is a single HTML file with no external dependencies, so it can be opened directly or hosted anywhere.

Was this page helpful?