Tencent consolidated nine monitoring tools into one graph-first platform over 8 million nodes
Tencent operates internet-scale cloud and infrastructure services where reliability, fast incident response and deep system visibility are non-negotiable. Its infrastructure monitoring stack now runs on SurrealDB, with document data and graph relationships in one engine.
Challenge
Before SurrealDB, Tencent's monitoring and analysis spanned a patchwork of nine systems: MySQL, Elasticsearch, VictoriaMetrics, MongoDB, Doris, Trino, RisingWave, Flink and Dgraph. Every investigation began with a meta-problem, choosing the right storage or compute engine before the question could be asked.
For analysts that choice was hard without deep familiarity with each engine's trade-offs. For the platform team the cost was plainer: more systems to operate, more upgrades and failure modes, more governance to enforce and a higher maintenance burden.
The most valuable workflows were graph-shaped. Process behaviour, parent and child derivations and traffic relationships had to be modelled and then queried fast during an incident. The data arrives as periodic process snapshots and real-time changelogs, from which Tencent builds a process tree, and that tree has to stay continuously updated and versioned so an engineer can reconstruct the state leading up to a fault.
Solution
Tencent adopted SurrealDB as a unified data layer for monitoring correlation and graph-driven fault analysis, consolidating the toolchain into one system that natively supports document data and graph relationships. The shape of the problem drives the query, not the choice of engine.
Processes and infrastructure entities are nodes and their relationships are edges, so engineers traverse multi-hop dependencies to understand blast radius and root cause. It is the same context-graph approach many teams use to build operational knowledge graphs.
Versioned data access mattered as much. Tencent validated the temporal requirements on SurrealKV, replaying the evolution of a process tree for point-in-time investigation, and runs production on SurrealDB with a distributed storage layer: 9 storage nodes and 6 compute nodes behind one operational experience.
Results
Nine tools became one
MySQL, Elasticsearch, MongoDB, Dgraph, Flink and four other tools were replaced with a single platform.
Production-grade scale
The cluster sustains more than 10,000 queries per second across the monitoring workload.
An 8 million node process graph
With continuous real-time updates as snapshots and changelogs arrive.
50 million edges, and a knowledge graph next
The edges across 8 million nodes are the foundation for a broader operational knowledge graph.
What is next
Tencent plans to extend the platform across more infrastructure domains, connecting services, processes and dependencies in one queryable model, with a focus on reconstructing historical states and building graph analytics over millions of nodes and edges.








