On this page
A production AI application usually runs a small zoo of stateful systems. Each has its own cluster, replicas, upgrade cadence, security model and on-call rotation:
| What the app needs | What teams usually run | What it holds |
|---|---|---|
| Event ingest, agent traces, change data | Kafka | Raw events |
| Keyword and hybrid search | Elasticsearch | Copy 2 of the documents |
| Semantic retrieval | Qdrant, Pinecone | Copy 3: embeddings and payloads |
| Knowledge graph, GraphRAG, agent memory | Neo4j | Copy 4: entities and relations |
| Product analytics, evals, cost dashboards | ClickHouse | Copy 5 |
| Agent runs that retry, wait and resume | Temporal, or a queue plus cron plus Postgres | Workflow state, in yet another database |
| The application's own data | Postgres, Firebase, Convex | Users, sessions, settings |
Three problems follow from that table. The same entity lives in four stores glued together by connectors, so when they drift, an agent retrieves stale or contradictory context. The retrieval glue lives in application code: one "hybrid GraphRAG" lookup is BM25 in one service, nearest neighbours in another, a graph hop in a third and a fusion step in Python, and nothing can plan it as a whole. And each store replicates two or three times on block storage and runs around the clock, sized for its peak.
Loam is our answer: one platform for the data side of an AI application, built on object storage and open formats, with compute that holds nothing it cannot rebuild. This post is the first in a series that walks through each part. It says what the platform is, why object storage is the foundation, where we deliberately step off it, and exactly what exists today.
What Loam is
Loam started as a retrieval engine and grew into what we call an AI-native cloud. It has three parts that share one binary, one tenancy model and one bucket, plus four that are proposed and not yet built.
Stateless roles: gateway, log, query, worker. Any node can be killed.
- The retrieval engine. Collections of documents with dense and sparse vectors, full text and filters, searched with one planned hybrid query. It speaks the Qdrant API, a subset of the Elasticsearch API, Arrow Flight SQL and its own REST API. Every write lands in a log on object storage first. Part 2 covers it.
- Loam Live, a reactive application database in the style of Convex, on TiKV: documents, indexes, transactions, queries that stay live and push new results, and a change feed into search. Part 3.
- Durable execution. The Resonate server, linked into the Loam binary, so the official Resonate SDKs get durable promises, retries, sleeps, schedules and human-in-the-loop gates from the same system that holds the agent's memory. Part 4.
- Functions billed on CPU time, a proposal: agent and backend functions that run next to the data and cost nothing while they wait on a model. Part 5.
- Jobs, a proposal: Celery, BullMQ, PySpark and Flink SQL workloads on Loam with a configuration change, behind one Rust API. Part 6.
- Postgres and MySQL, a proposal: Neon and WeSQL run beside Loam on the same object store, with a branch per agent workspace. Part 7.
- Self-hosting with GitOps, a proposal for the whole stack, and the line between what is open source and what only the managed cloud runs. Part 8.
Why object storage
The retrieval engine's first design rule is that object storage is the only durable source of truth, and compute is stateless. We wrote about the cost side in Why a retrieval engine should live on object storage. The short version:
- Cost. Replicated block storage costs roughly $0.16 to $0.24 per effective GB-month (EBS gp3 at about $0.08, times two or three replicas). S3 Standard is about $0.023, with durability handled by the provider.
- Idle data is the common case. AI applications create an index per user, per workspace, per repository or per agent. Most are idle most of the time. In Loam, a cold namespace is a set of immutable objects. It costs its storage and nothing else until someone queries it.
- Open formats. Documents and vectors are a Lance dataset, full text is Tantivy split files, and analytics tables will be Apache Iceberg. Other engines can read them without Loam. If you leave, your data is still readable where it sits.
- Nodes you can lose. Each node holds caches and derived structures only. Killing one loses cache warmth, not data. Scaling is adding or removing stateless processes.
What made this practical is recent. S3 added conditional writes (If-None-Match and If-Match on PUT) in 2024, and GCS and Azure have equivalents, so a single object can be a compare-and-swap commit point. The Rust data stack (DataFusion, Arrow, Lance, Tantivy, object_store, foyer) matured enough to build a serving layer on. And the architecture is proven in production by closed products: turbopuffer for search, WarpStream for streams, ClickHouse Cloud and StarRocks for analytics. No open-source project had applied it to vector, text and graph retrieval in one engine.
The price you pay
Object storage is slow to write to compared with a local disk. A write in Loam is acknowledged only after it is durable in the bucket: one PUT of a batched WAL object plus one metastore commit. That is tens of milliseconds, not microseconds. In exchange, an acknowledged write survives the loss of any node, and a query can read it immediately: every write returns a consistency token, and a read that carries it merges the not-yet-indexed tail of the log, so read-your-writes never waits for indexing.
Reads from a cold bucket pay object-storage round trips too. The hot tier (RAM and NVMe caches, and prebuilt HNSW graphs for hot collections) exists to make the warm path fast. Correctness never depends on it: results are identical with the hot tier on or off, and the test suite checks that.
Where Loam steps off the bucket
Not everything belongs on object storage, and we would rather say so than hide it.
- Metadata. Offsets, leases and manifest pointers change hundreds of times a second. S3 conditional PUTs cannot sustain that, so metadata lives in a metastore: embedded Raft (openraft with a redb log) by default, or TiKV in clusters. Manifest bodies stay immutable objects in the bucket; only pointers live in the metastore.
- Application data in Loam Live. An application database needs millisecond transactions over mutable rows. That is what TiKV does, with Raft replication on local disks. So Live's source of truth is TiKV, not the bucket. The Live role itself is stateless, and the design makes continuous log backup of every Live keyspace to object storage mandatory before a cluster serves traffic. That backup is designed, not built.
- Durable execution and jobs state. Workflow promises and job leases are small, hot, transactional records. They live in SQLite on a single node, and in TiKV (or, today, TiDB) in clusters. Large payloads and results go to the bucket.
The rule we hold: the retrieval engine keeps the bucket as its only source of truth, and anything that needs OLTP transactions runs on TiKV beside it, never inside it.
One tenancy model, one binary
Everything is organised by namespace. Quotas, encryption keys, cache affinity, routing and billing are per namespace. A namespace can hold collections and streams, one Live app and its durable workflows. The operon dev command (the binary is renamed loam once the first milestone ships) runs every role in one process on a laptop; a cluster runs the same binary with roles split across pools.
The platform is open source under Apache-2.0. Its dependencies are held to the same bar: no AGPL, BSL, SSPL or ELv2 code is linked into the engine. Services with other licenses, such as WeSQL (GPL-2.0), are only ever run unmodified as separate processes, and we say so where it matters. Part 8 explains what the managed cloud adds on top.
Where this stands
Loam is being built in the open, and much of this series describes design work. Every post ends with a table like this one. Available means it is on the engine's main branch with its tests passing. In progress means part of it is on main or on a branch under review. Planned means designed or proposed, with no code yet.
| Part | Status | What exists |
|---|---|---|
| Log, metastore, workers, links, crash gates | Available | The foundation: WAL on object storage, openraft metastore, exactly-once links, kill -9 tests at every commit step |
| Collections and hybrid search | Available | Lance and Tantivy under one manifest, vector, BM25 and sparse retrieval, fusion, the tail merge |
| Hot tier and cluster routing | Available | RAM and NVMe cache, HNSW artifacts, affinity routing, write backpressure |
| Qdrant API, Flight SQL | Available | Qdrant REST and gRPC, Flight SQL with bulk ingest |
| Elasticsearch subset | In progress | Most of it on main |
| SDKs and MCP server | In progress | Built on branches, landing |
| Loam Live and the TiKV metastore | In progress | Transactions, the commit journal, subscriptions and the TiKV metastore on main |
| Durable execution | In progress | Resonate embedded with SQLite and TiDB stores on main, behind a build flag |
| Graph expansion, Iceberg tables, stream API, Kafka gateway | Planned | Designed |
| Functions, jobs, Postgres and MySQL, GitOps | Planned | Proposals under review |
The roadmap tracks the same states. The next post starts with the part that exists most completely: the engine.