HyperCrux, an open-source database that reached version 0.1 in October 2026, keeps records, the links between them and their vector embeddings in a single SQLite file. An application can reach any record four ways: by its key, by SQL, by following its links, or by asking which records sit closest to a question in vector space. There’s no server. It ships as a Go library and as one small binary that any language can call, while the rules that keep the four views in step live inside the file as SQLite triggers.
A release this new has no adoption record to judge. The design does have a market. That market is the subject here: who wants one small database for lookups, joins, link walks and similarity search over the same data, and where that demand runs out.
The short answer is AI agents, plus AI that runs locally. Both create small databases in very large numbers. Neither wants the same facts kept in three systems.
Three Copies of the Same Facts
When an ordinary app gets AI search, its data tends to split. Records stay in the relational database. Embeddings move to a vector store so the app can find documents that resemble a question. Relationships, such as who owns which file or which paper cites which, end up in a graph database or in join tables. Glue code copies every change from one system to the others; when it misses one, the answers stop agreeing. A search returns a document deleted an hour ago. A walk follows a link to a record that no longer exists.
The failure is quiet, which makes it costly. Users meet it as a support bot quoting a withdrawn article, or an assistant recalling something it was told to drop. Under the GDPR’s right to erasure it also becomes a legal problem. An embedding of personal text counts as personal data whenever it can be tied back to the person. Researchers have also shown that embeddings can be partly turned back into the text they came from. Deleting the source row while its vector lingers in another system leaves the job half done.
HyperCrux keeps one copy. A record is a row: its key names it, its fields are columns, it can carry a vector and it can link to other records. Deleting the record removes its links in the same transaction. Because the rules sit in the file (the triggers fire whichever program makes the change), a Python script writing plain SQL is held to them just as the Go library is. Before release, a writing process was killed at random moments 200 times without leaving a single mismatched record, as the write-up on how HyperCrux was tested details.
The 2026 Landscape: Big Engines and Small Files
Two camps are working on this problem from opposite ends. The first builds one large engine for every data model. SurrealDB released version 3.0 in February 2026 alongside a $23 million extension to its Series A, taking total funding to $44 million. It now markets the product as memory for AI agents, with relational, document, graph, vector and key-value data behind one query language. HelixDB, a Y Combinator company, joins graph and vector storage in a Rust engine. Its founders describe the pain in the same terms HyperCrux does: separate vector and graph databases held together by custom sync code.
The second camp starts from SQLite, with the idea that a database should cost no more to create than a file. Turso rewrote SQLite in Rust with built-in vector search, then built a cloud where one server can run millions of small databases; on 2 October 2026, Supabase agreed to buy it. The sqlite-vec extension, a Mozilla Builders project, adds vector search to ordinary SQLite. Kuzu, an embedded graph database that grew out of research at the University of Waterloo, was bought by Apple in October 2025; its code continues as a community fork. Through 2026, several smaller open-source engines appeared that pair SQLite-style SQL with vector search; one adds graph queries too.
Postgres, the database most developers use, covers part of the ground through extensions: pgvector for embeddings, with graphs left to other extensions or to join tables. Native graph queries based on the SQL/PGQ standard were committed for Postgres 19 in March 2026, then removed in September, before release, over readiness concerns. For now the mainstream route to SQL plus graph is still a workaround.
| Product | Data models | How it runs | Status in 2026 |
|---|---|---|---|
| SurrealDB 3.0 | Relational, document, graph, vector, key-value and more | Server or embedded, own query language | General release in February, aimed at agent memory |
| HelixDB | Graph and vector | Server and managed cloud, own query language | Y Combinator-backed, AGPL licence |
| Turso | SQLite-compatible SQL with vector search | Embedded engine plus cloud | Being acquired by Supabase |
| sqlite-vec | Vector search for SQLite | SQLite extension | Still before version 1.0 |
| Kuzu, now LadybugDB | Graph (Cypher) with vector search | Embedded | Company bought by Apple; community fork |
| Postgres with pgvector | SQL with vector search | Server | Native graph queries dropped from version 19 |
| HyperCrux | Keys, SQL, typed links, vectors | Go library or one binary; plain SQLite file | Version 0.1, Apache 2.0, October |
HyperCrux is the smallest entry on this map: all four handles in a file that any SQLite tool can open. Next to the big engines it gives up scale and a rich query language. Next to the SQLite-based tools it adds what they lack, which is typed links plus one consistency rule covering key, row, link and vector.
Level of Interest: From Vector Stores to Memory
In 2023 the money went to standalone vector databases. That category is shrinking into a feature. Most major databases now offer vector search, and Pinecone, the company that defined the category, has explored a sale. Independents that still raise money, such as Qdrant with its $50 million Series B in March 2026, compete on scale and cost.
New money is flowing to memory. Mem0 raised $24 million in October 2025; AWS chose it as the memory provider for its agent SDK. Cognee raised a seed round in February 2026 for graph-structured agent memory, with an engine for edge devices on its plan. Developer pull is stronger still. MemPalace, an open-source memory project, gathered more than 19,500 GitHub stars in its first week in April 2026 and passed 56,000 within months; Mem0’s repository has passed 59,000. Database researchers have started to treat long-term agent memory as a workload of its own, listing forgetting next to retrieval as a core operation.
The clearest signal is in acquisitions. Databricks agreed to pay about $1 billion for Neon in May 2025, when more than 80% of new databases on Neon were being created by AI agents. Apple bought Kuzu five months later. Supabase launches more than a million databases a week, about 70% of them created by agents or AI tools. On 2 October 2026 it raised $150 million and agreed to buy Turso. The buyers differ. The bet is the same: agents will create databases far faster than people ever did, and most of those databases will be small.
HyperCrux enters at the small end of that bet.
Demand Drivers
Agent memory is the largest driver. An assistant that remembers needs rows for facts about people, links for the relationships between them, vectors for recall by meaning and keys for fast lookups. All of it has to change in one step. When a user corrects a fact, the old embedding must stop matching. When a user asks to be forgotten, the person and every fact and vector derived from them have to go together. Teams that build this on separate stores end up writing the sync code themselves. The post on an AI assistant’s memory that forgets properly shows the same job done in one transaction.
Agents also change how many databases exist. A coding agent spins up a database for each prototype; an assistant keeps one for each user. Platforms now compete on how cheaply an agent can create a database. HyperCrux already is a file. Creating one takes a name, and backing it up takes a copy.
Retrieval is moving past vector-only search. Vector search finds passages that sound like the question, but it can’t tell that one clause amends another or that two incidents hit the same service. Hybrid retrieval, which uses vector search to find entry points and graph steps to collect what’s connected, is becoming the default enterprise pattern. In application work the graph step is usually short, one to three hops from the record the vector search found. That’s a different job from graph analytics across billions of edges.
Local AI moves the database onto the device or into the office. Embedding models small enough to run offline, such as Google’s 308-million-parameter EmbeddingGemma, let a laptop or phone embed its own documents. Law firms, clinics, recruiters and accountants all have reasons to keep client data on their own machines. One file that never leaves the building is easy to explain to a client or an auditor.
Regulation sharpens the point. Hiring tools sit in the EU AI Act’s high-risk category; the Digital Omnibus, approved in June 2026, moved the related obligations from August 2026 to December 2027 without removing them. GDPR erasure applies now. A recruiting agency that matches candidates by embedding has to show that a deleted candidate is gone from search as well as from the table. That case is worked through in matching candidates to jobs and deleting them cleanly.
Technology Drivers
SQLite is the base layer. It’s the most widely deployed database engine in the world, used by more than a third of developers. Its file format is stable enough to sit on the US Library of Congress list of recommended formats for datasets. Building on it means every mainstream language and every SQLite browser can already open a HyperCrux file. The trigger-based design turns that reach into a guarantee, since third-party programs can write to the file without breaking it.
Hardware has made brute force respectable. Small embedding models produce vectors of a few hundred values (384 is common), which a modern CPU compares quickly. On a two-core cloud machine, HyperCrux finds the 10 nearest among 10,000 vectors of 384 values in about 42 milliseconds, and among 100,000 in 0.41 seconds. Exact search needs no index tuning and gives up no recall. A SQL filter applied before the comparison costs nothing in accuracy either. For the long tail of small datasets, that’s simpler than any approximate index.
Agents now reach tools through a standard interface. The Model Context Protocol passed 97 million monthly SDK downloads in March 2026, with more than 10,000 public servers. A database that sits in a file next to the agent, with a command line that reads and writes JSON, fits that world without a network service.
Graph support in mainstream SQL engines hasn’t settled. SQL/PGQ became an ISO standard in 2023, yet Postgres removed it from version 19 before release. Until Postgres or SQLite supports graph queries natively, applications that need typed links and short walks will reach for small tools that provide them. HyperCrux walks one link out from a record in about 43 microseconds on a graph of 100,000 records with five links each.
Verticals and Uses That Fit
A use fits HyperCrux best when four conditions hold. Each table stays under roughly 100,000 vectors. The application needs at least two handles at once, such as similarity filtered by SQL or a link walk ranked by similarity. Deletions have to be complete. And the data can live on one machine; where privacy matters, that’s an advantage in itself. The eight cases published with the release meet that test, and they fall into five kinds of buyer.
Teams building AI products are the largest group. Assistant and agent memory is the core use. Customer support comes next: a help-centre bot ranks articles by meaning while SQL filters out retired ones. A stale index is exactly what makes a bot quote an article withdrawn last week. The case is laid out in a support bot that never quotes a retired article.
Professional firms often hold sensitive data with little IT staff. A law firm’s clause library pairs similar-clause search with links from each clause to its precedents and clients. It also has to stay in the office (see a law firm’s clause library). A recruiting agency needs candidate matching by embedding, plus deletion a regulator would accept. For both, the alternative is a hosted vector service they would have to justify to their own clients.
Operations teams search past incidents by similarity, then follow links to the services and changes involved. Incident histories are small. A local file still answers when the wiki is down at 3 a.m., the scenario in find the incident that looks like this one.
Individuals with large personal collections form a fourth group. A PhD student’s literature review links papers by citation, searches abstracts by meaning and filters by year with SQL. A photographer’s archive searched by image embeddings finds the shot its owner half remembers. Neither user will run a server.
Small businesses are the fifth. A bike shop’s recommendations touch all four handles: similar products by vector, compatible parts by link, stock by SQL, items by key. A small retailer won’t run three databases. One file inside its shop software is plausible.
The same test points to uses the release posts don’t cover. Desktop and mobile apps with on-device AI can give each user a file of their own. SaaS products can give each customer one, so a tenant’s data can be deleted in one move.
Where the Demand Stops, and What Could Move It
The limits are published in the project’s README. Search compares every vector, so beyond about 100,000 vectors per table an approximate index wins; that business belongs to pgvector and the dedicated vector databases. Deep traversals across large, dense graphs belong in a graph database. SQLite’s locking doesn’t work over network file systems. It also allows one writer at a time. A file shared across machines, or one under heavy concurrent writes, needs a server database. Embedding the Go library needs cgo and a C compiler, which the binary avoids.
Those limits also mark out the market. HyperCrux is built for the small databases that agents and local apps create by the million. The deals of 2025 and 2026 show the large platforms expect that kind of database to multiply. Four developments will show how fast:
- How Supabase uses Turso. If a large platform makes SQLite-shaped databases the default home for agent data, the one-file model goes mainstream.
- What Apple does with Kuzu. Graph databases on phones and Macs would make on-device graph memory ordinary.
- Whether SQL/PGQ returns in Postgres 20, which decides how long Postgres users keep relying on workarounds for graph queries.
- Whether sqlite-vec reaches version 1.0 and agent memory frameworks add local backends. Either would widen the pool of developers who expect a capable database inside a file.