Skip to content

Kreuzberg & SurrealDB: unstructured docs to hybrid search

Release

Apr 30, 20262 min read

Ignacio Paz

Ignacio Paz

Show all postsKreuzberg & SurrealDB: unstructured docs to hybrid search

We’re excited to share a new partner integration: kreuzberg-surrealdb, a connector that bridges the Kreuzberg document intelligence framework directly into SurrealDB. This integration was created by the Kreuzberg team and we are excited to have this functionality available now in SurrealDB.

Kreuzberg extracts, chunks, and generates embeddings from 88+ document formats, while SurrealDB provides a multi-model database for AI applications, combining documents, graphs, vectors, and full-text search in a single system.

Together, they make it easy to build document search and RAG pipelines.

Our newsletter
Get tutorials, AI agent recipes, webinars, and early product updates in your inbox every two weeks

Three-stage pipeline: unstructured documents (PDFs, images, docs) go through
extraction into Kreuzberg, which handles ingestion, chunking and embedding,
then into SurrealDB for full-text, vector (HNSW) and hybrid (RRF) search.

kreuzberg-surrealdb handles the full ingestion workflow:

  • Automatic schema setup

  • Content deduplication using SHA-256 hashing

  • Storage and indexing in SurrealDB

  • Documents ready for search immediately after ingest

The integration supports two modes:

  • DocumentConnector:indexes full documents for BM25 keyword search.

  • DocumentPipeline:chunks documents, generates embeddings, and enables semantic and hybrid search using HNSW vector indexes and Reciprocal Rank Fusion.

Building document search systems often requires combining multiple tools for extraction, chunking, embeddings, and storage.

With kreuzberg-surrealdb, the entire workflow runs through a single integration - no schema boilerplate, no duplicate ingestion, and built-in support for keyword, semantic, and hybrid search.

See how to get started in SurrealDB Docs: Kreuzberg Integration, and check out our example of How to build a knowledge graph for AI with SurrealDB and Kreuzberg.

Related posts

Our newsletter

Get tutorials, AI agent recipes, webinars, and early product updates in your inbox every two weeks

SurrealDB

The context and memory layer for AI agents

Database. Graphs, vectors, documents and relational data in one engine, in a single ACID transaction.
Agent Memory. Connects and retrieves context wherever your data lives, every fact carrying its source.
Cloud. Fully managed, in the cloud provider and region you choose.

Explore with AI

Copyright © 2026 SurrealDB Ltd. Registered in England and Wales. Company no. 13615201

Registered address: 3rd Floor 1 Ashley Road, Altrincham, Cheshire, WA14 2DT, United Kingdom

Trading address: Huckletree Oxford Circus, 213 Oxford Street, London, W1D 2LG, United Kingdom