Skip to content

Overview

Scaling

This page explains when to change the resources behind an instance, and what to expect from each kind of change.

What you can change depends on the plan. Start instances scale vertically on one node. Scale clusters scale vertically and horizontally.

The steps for making a change are in Configure an instance.

StartScale
Compute upA larger instance type on one nodeA larger node size, or more nodes
Compute downA smaller instance typeA smaller node size, or fewer nodes
StorageIncrease onlyIncrease only
Fault tolerance while resizingA brief reconnect while the node restartsThe cluster stays available. Capacity dips as each node restarts

Storage can be increased but never decreased on either plan, and you can increase the disk size once every six hours. Over-provisioning storage is therefore a one-way decision. Start close to what the dataset needs and grow it as the dataset does.

The signals worth acting on are:

  • CPU pegged: the instance sits near its instance-type limit under normal traffic, not only at peaks.

  • Memory pressure: the working set no longer fits, so the engine reads from disk more often.

  • Storage growth: consumption trends towards the provisioned limit. Act well before the limit, because a full disk stops writes.

  • Connection saturation: clients queue or are refused.

  • Latency objectives breached: queries that used to meet them no longer do.

Read those signals from the instance dashboard, metrics, and logs.

Rule out query-level causes first. A missing index or a full-table scan is cheaper to fix than a larger instance, and a resize hides the problem rather than removing it.

Schedule resizes outside peak traffic where you can. Some changes cause a brief reconnect, so clients need retry logic in any case.

A Start instance has one node, so growth means a larger instance type and a larger disk. A larger type gives more vCPU, more memory, and more I/O.

A Free instance moves to a paid instance type in place, in the same way as any other type change. It does not need a new deployment or an export.

The storage range of a Start instance depends on the memory of its instance type. A paid type starts with 4 GB of storage for each GB of memory, and its storage can be raised to 8 GB for each GB of memory. The Free type has a fixed 1 GB disk.

Instance typeMemoryStarting storageMaximum storage
Free512 MB1 GB1 GB
Burstable small1 GB4 GB8 GB
Burstable medium2 GB8 GB16 GB
Burstable large4 GB16 GB32 GB
General purpose medium4 GB16 GB32 GB
General purpose large8 GB32 GB64 GB
General purpose xlarge16 GB64 GB128 GB
General purpose 2xlarge32 GB128 GB256 GB
General purpose 4xlarge64 GB256 GB512 GB

To grow past the maximum of the current type, move to a larger type first. For example, 50 GB of storage needs at least General purpose large, whose maximum is 64 GB.

A type change sets the storage to the starting storage of the new type, or keeps the current size if that is larger. The type change also counts as a storage change, so the six-hour wait starts again before the storage can be raised on the new type. Where the starting storage of a type already covers what you need, choosing that type avoids the wait.

Move to Scale instead when a single node is no longer the right shape:

  • The workload must survive the loss of a node.

  • Query throughput needs to spread across nodes.

  • You have reached the largest available Start instance type.

Moving from Start to Scale is not an in-place upgrade. Deploy a Scale instance, then restore a backup or import an export into it.

A Scale cluster grows in two directions. Vertically, each node moves to a larger size. Horizontally, you add nodes, so a three-node cluster becomes four, then five, and so on. The storage layer handles replication, consensus, and distributed transactions across whatever the node count becomes.

Note

An even node count does not improve fault tolerance. A cluster of N nodes survives the same number of failures as one of N−1 nodes. The fast commit path also needs every node to agree, so one slow node forces the slower path. Odd sizes of three, five, or seven give the best ratio of resilience to cost.

After you scale out, check two things. Confirm that replication lag stays within what the application tolerates. Then check per-node balance, because an unbalanced cluster means the added nodes are not taking their share of the load.

See Architecture for the topologies, and High availability for what each one survives.

Compare the same metrics before and after the change. A resize that shifts the bottleneck rather than removing it is common: CPU headroom appears, and disk I/O becomes the new limit.

If utilisation stays low after a large resize, step back down. Instance type is billed on what is provisioned, not on what is used, so an oversized instance is a standing cost with no benefit.

surrealctl instance metrics and surrealctl instance update let you read the signals and apply the change from a terminal. See surrealctl instances.

Was this page helpful?