Posted 11 min read

SurrealDB: the all-in-one database we chose over PostgreSQL

The perfect database is usually several databases stitched together. We weighed PostgreSQL and its extensions, and chose one engine that is all of them.

EngineeringDatabasesAI

The words "The unified data layer for AI" in white on violet, beside the SurrealDB mark
On this page
  1. What is a multi-model database?
  2. Why not PostgreSQL with extensions?
  3. Every extension is a system of its own
  4. What extensions cannot absorb
  5. Scaling and availability are add-ons
  6. How SurrealDB works as a web database
  7. Many companies, many apps
  8. A group of companies is a graph
  9. Built for AI context graphs
  10. Performance: Rust and one engine
  11. Scalability and high availability
  12. Rust, and source available
  13. Held to account: a schema the database enforces
  14. What it costs us
  15. The foundation we want

Foretag builds and operates companies in very different industries, from productivity software to travel and retail. Each company runs its own apps, and each app keeps its own data. So the database is not one decision among many. It is the layer every company stands on, and the one we are least willing to get wrong.

We chose SurrealDB. It is a multi-model database written in Rust: relational, document, graph, key-value, vector, full text, geospatial and live data in one engine, queried in one language and secured by permissions the database enforces itself. Its source is public, and the same queries run from an embedded engine in a test to a distributed cluster in production. We looked hard at PostgreSQL with extensions first. What follows is what we compared, why one engine won, and what the choice costs us.

What is a multi-model database?#

A multi-model database stores and queries more than one shape of data natively, in one engine, instead of asking you to run a separate system for each. Most applications need several shapes at once: records to store, relationships to follow, text to search, places to map, meaning to compare and changes to stream. The usual answer is a system for each shape, and a pipeline between every pair.

In SurrealDB, every shape belongs to the same data. One record can hold nested documents, link to other records, sit at either end of a graph edge, carry a vector and a location, and stream to a client as it changes, all under one schema and one set of permissions.

Records live in tables that can be as strict or as flexible as their schema says, and each is addressed by an ID that is a direct key. A field can link to another record, and following the link is a dot, not a join.

CREATE company:whole CONTENT {
	name: 'Whole',
	sectors: ['Travel', 'Retail'],
};

CREATE store CONTENT { name: 'Flagship', owner: company:whole };

SELECT name, owner.name AS company FROM store;

Why not PostgreSQL with extensions?#

We did not arrive at SurrealDB by default. PostgreSQL was the serious alternative, and it deserves its reputation: three decades in production, a superb query planner, strong transactional guarantees, and an extension ecosystem that can turn it into almost anything. So we sketched what we needed as a Postgres stack, next to the same needs in SurrealDB:

NeedPostgreSQLSurrealDB
Relational and documentsTables and JSONBTables, nested objects and record links
GraphApache AGE (Cypher inside SQL), or recursive CTEsEdge tables, RELATE and arrow traversal
Vector searchpgvector (HNSW or IVFFlat)HNSW index
Full texttsvector with ts_rank, or ParadeDB for BM25Full text index with BM25
GeospatialPostGISGeometry types and geo:: functions
Live updatesLISTEN/NOTIFY or logical decoding, and a delivery serviceLIVE SELECT
Row level securityBuilt inTable and field permissions
Browser clientsAn API layer, such as PostgRESTDirect, over WebSocket or HTTP
Write scalingSharding, such as CitusWrite nodes in SurrealDS
Automatic failoverPatroni, or a managed serviceQuorum writes across zones, no leader

Every entry in that middle column is good software. The trouble is the sum.

Every extension is a system of its own#

An extension brings its own types, operators, index methods, tuning parameters and release cycle. pgvector, PostGIS and Apache AGE are maintained by different people on different schedules, and a major Postgres upgrade waits until every one of them supports the new version. Not every managed provider offers every extension, so the extensions you choose narrow the hosts you can use, and the other way round.

Each one also speaks its own dialect inside SQL: Cypher passed as a string to AGE's cypher() function, PostGIS's spatial functions, pgvector's distance operators, and the tsvector and tsquery of built in full text search. Postgres can combine all of them in a single statement, and that is a real strength. But it is still several type systems to learn, index and tune, and several things that can change underneath you.

What extensions cannot absorb#

Some needs do not fit inside Postgres at all. LISTEN and NOTIFY are not durable, carry payloads of under 8,000 bytes by default, and are not filtered by row level security, so live updates usually mean logical decoding and a service to deliver the changes. Browsers cannot speak the Postgres wire protocol, so direct client access means an API layer such as PostgREST in front of the database. And because Postgres gives every connection an operating system process of its own, a large number of clients means a pooler such as PgBouncer.

Each of those services is another place that has to know who may see what. Row level security protects the tables, but a delivery service or an API layer either needs the rules again or has to be trusted to pass them through faithfully. That is where permissions drift.

Scaling and availability are add-ons#

Postgres scales reads well with streaming replicas, but every write goes to a single primary, and scaling writes out means sharding, with Citus or in the application. Automatic failover is not built in either: it comes from Patroni, a similar tool, or a managed service. All of them are proven, and all of them are more systems to run, upgrade and understand.

None of this is a case against Postgres. We still love it, and we run it for the many third party applications that support nothing else, where it is exactly the right choice. For a single relational application that needs one or two of these capabilities, it is hard to beat.

The applications we build ourselves are a different case: many companies and many apps, most of which need most of these shapes at once. For those, we choose SurrealDB. Put side by side, the difference is plain:

8systems to run and secure

Your app
  • GraphGraph database
  • RelationalRelational database
  • DocumentDocument database
  • Key-valueKey-value store
  • VectorVector database
  • Full textSearch engine
  • GeospatialGeospatial database
  • Live queriesMessage broker

Eight systems, each with its own copy of the data, its own interface and its own permissions.

SurrealDB: one copy, one query language, one set of permissions.

Fewer systems is not only less to run. It means fewer copies of the data to drift out of step, and fewer places to get security wrong.

How SurrealDB works as a web database#

We use SurrealDB as a web database. Once someone has signed in, the app connects to its database directly, over WebSocket or HTTP, as that person, with a short lived token the database verifies for itself. There is no server in the middle rewriting queries, and so no server to get permissions wrong.

The database decides who sees what.

Every table carries its own permissions, written next to the data they protect, and a field can carry permissions of its own:

DEFINE TABLE invoice SCHEMAFULL
	PERMISSIONS
		FOR SELECT WHERE company IN $auth.companies
		FOR CREATE, UPDATE WHERE company IN $auth.companies AND $auth.role = 'finance'
		FOR DELETE NONE;

DEFINE FIELD notes ON invoice TYPE option<string>
	PERMISSIONS FOR SELECT WHERE $auth.role = 'finance';

Permissions are evaluated when data is read, not written into a token when it is issued, so removing someone from a company takes effect on their very next query. They apply to live queries and graph traversals as well: a record the user may not see is left out wherever a query reaches it. And an app that forgets to check cannot get around them.

Many companies, many apps#

SurrealDB organises data as namespaces that hold databases, and each level can have its own users and access rules. That hierarchy suits a group of companies. Ours share engineers, practices and tools, not data: each company, and often each app within it, has a database of its own, so one company's records never sit beside another's.

Share the tools, never the data.

Because every one of those databases is SurrealDB, a new company starts with the same schema patterns, the same tests and the same tooling as the last, without sharing anyone else's data.

A group of companies is a graph#

Foretag is a group: a parent, operating companies, and subsidiaries below those. Plenty of businesses have the same shape, and most software flattens it into a single account. In SurrealDB the structure is simply what it is. A subsidiary is a relation from one company to another, with fields and rules of its own:

DEFINE TABLE owns SCHEMAFULL TYPE RELATION FROM company TO company;
DEFINE FIELD since ON owns TYPE datetime;
DEFINE FIELD stake ON owns TYPE number ASSERT $value > 0 AND $value <= 100;

People, teams and the companies they work for are related in the same way, so a question that spans the group is one traversal rather than a chain of joins. Everyone who works at any of our subsidiaries, for example:

SELECT ->owns->company<-works_at<-person.name AS people
FROM company:foretag;

Built for AI context graphs#

An AI system is only as useful as the context it is given, and context is a graph. The documents that answer a question were written by people, who belong to organisations, who work on projects, and every one of those carries rules about who may see it. Retrieval that follows those connections, often called GraphRAG, needs vectors, a graph and permissions at the same time.

Built from separate systems, that means a vector store to find relevant text, a graph database to follow the connections, and application code to filter out what the user may not see, stitched together on every request. In SurrealDB it is one query, and the permissions come with the data. Step through it:

QuestionDocumentsPeopleCompanies and projectsQuestionWhere should we open next?DocumentFootfall by cityDocumentAvailable leasesDocumentTravel trendsDocumentSummer party planPersonAmiraPersonTheoCompanyRetail companyProjectExpansion planEverything lit is the context for the model

A question arrives

Someone asks the assistant where to open the next store. The answer is somewhere in the company’s documents, and in who wrote them.

SELECT	title,	->written_by->person.name AS authors,	->written_by->person->works_at->company.name AS companies,	->about->project.name AS projectsFROM documentWHERE embedding <|3, 40|> $question;

None of this needs a new system. The records, relationships and permissions already live in the database, and vectors are one more index on the same tables.

Performance: Rust and one engine#

SurrealDB is written in Rust, and for a database that is more than a matter of taste. Rust gives memory safety without a garbage collector: whole classes of memory bugs are ruled out at compile time in safe code, and there are no collection pauses to surface in tail latency. Connections are served as lightweight asynchronous tasks rather than an operating system process each, which matters when every browser is a client.

Most of the gain, though, comes from architecture rather than language:

  • Fewer hops. A query that finds documents by meaning, follows their relationships and applies permissions runs in one engine, in one round trip, rather than across three services.
  • Nothing to keep in sync. Indexes are maintained by the database itself as records change, so there are no dual writes to get wrong and no pipeline to fall behind.
  • Direct addressing. Record IDs are keys, and record links and graph edges point straight at the records they connect, so following one is a lookup rather than a join worked out at query time.
  • No network when embedded. The same engine runs in process, in memory or on disk, and in the browser through WebAssembly. That is how our tests run, and how an app can keep working at the edge or offline.

Benchmarks depend on the workload, and the only ones worth trusting are those run on your own data. What the architecture guarantees is that the work a stitched stack spends moving data between its parts is simply not there.

Scalability and high availability#

SurrealDB separates query processing from storage, so one engine runs in very different shapes:

DeploymentStorageSuited to
EmbeddedIn memory or on disk, in process, including the browserTests, edge and offline apps
Single nodeRocksDB on diskMost apps, most of the time
DistributedSurrealDSHorizontal scale and high availability

Distributed, SurrealDB runs on SurrealDS, its own distributed storage layer, available on SurrealDB Cloud and with SurrealDB Enterprise. Compute is separate from storage, so query capacity and storage capacity scale independently, and capacity grows by adding nodes rather than by buying a bigger machine. Every write node can accept and coordinate transactions, and a transaction commits once a quorum acknowledges it, so there is no elected leader to fail over from. Clusters span availability zones and are designed to lose a node, or a whole zone, without dropping writes or losing consistency.

# A single node, on disk
surreal start rocksdb://data

The queries do not change. An app written against an embedded engine in its tests runs unmodified on a single node, and on SurrealDS when it outgrows one.

Rust, and source available#

Rust is also what our engineers reach for wherever correctness and performance decide the outcome, so SurrealDB is built the way we build.

Its source is public on GitHub under the Business Source License 1.1. That is a source available licence rather than an open source one, and it is worth being precise about what it means. We can read every line of the engine, build it ourselves and run it in production for our own products. What it rules out is offering SurrealDB itself to third parties as a database service. And each version converts to the Apache 2.0 licence four years after its release at the latest: for SurrealDB 3.0, on 1 January 2030.

For the layer everything else stands on, that matters. A database we can read is one we can question, debug and understand, and our tests start it in memory from the same binary we run everywhere else. It also fits how we think about independence. Our manifesto asks us to stand on our own, and a single engine we understand end to end is easier to own than a stack of services we only rent.

Held to account: a schema the database enforces#

A database we depend on has to be one we can check. Our tables are schemafull: fields have types, assertions and defaults, and events and functions carry the rules that span more than one record. A value that breaks the schema is rejected by the database itself, whichever code path tried to write it.

Keeping the rules there is one of the strongest reasons we chose SurrealDB. Permissions, assertions and events are written once, in SurrealQL, next to the data they govern, instead of being repeated in every app that touches it. They are reviewed in one place and enforced for every client alike.

The schema is tested the way code is. Every change runs our database tests, and we test ahead of the version we run, so a change in behaviour reaches us as a failing test rather than as an incident.

What it costs us#

No choice is free, and we have written down what this one costs:

  • A younger query planner. The Postgres planner has decades of tuning behind it. We read EXPLAIN on the queries that matter, and design record IDs and indexes deliberately rather than counting on the planner to rescue a poor model.
  • A different way to model. Thinking in records and relations, rather than tables and joins, takes practice, and a team new to it needs time to learn the patterns.
  • A fast moving system. Major versions do change syntax and behaviour, so we pin the version we run and test ahead of it.

We would rather pay these costs than spread one product's rules across eight systems.

The foundation we want#

A database is the part of a company that has to outlast every product decision made on top of it. SurrealDB gives each Foretag company, and each app within it, its records, relationships, search, places, vectors and live data in one engine and one language, secured by the database itself, with a path from a single process to a distributed cluster that needs no rewrite.

PostgreSQL with extensions could have taken us a long way. One engine takes us further, with less to run and less to get wrong. That is the foundation we want under everything we build.

Written by

  • CB

    Chiru Boggavarapu

    Founder & Chief Executive