Storage and the hot tier
What lives in the bucket, and what compute nodes keep to make it fast.
Every object in Operon has two tiers. The durable tier is an open format in object storage and the only source of truth. The hot tier is derived, node-local and rebuildable. Losing the hot tier changes latency, never correctness.
The bucket
Everything at rest lives under one prefix in one bucket:
s3://<bucket>/<cluster_prefix>/
meta/snapshots/… # metastore snapshots
wal/<class>/<node_id>/<ulid>.wal # WAL objects, many partitions and namespaces each
ns/<namespace_id>/
streams/<stream_id>/<partition>/<base_offset>-<ulid>.seg
collections/<collection_id>/
lance/… # Lance dataset
text/splits/<ulid>.split # Tantivy splits
manifests/<version>.pb # immutable collection manifests
hot/… # optional prebuilt hot-tier artifacts (HNSW)
graphs/<graph_id>/
idmap/… adj/… manifests/…
pk/<object_id>/… # primary key → row location
durable/wf/<origin> # one document per workflow origin
warehouse/<namespace_id>/<table_id>/ # Iceberg tables, catalogued in LakekeeperAll data objects are immutable and named by ULID or version. Only the metastore's pointers and the Iceberg catalog's pointers move. Operon relies on conditional writes (If-None-Match, If-Match), which S3, GCS and Azure all support.
The hot tier
| Object | Durable tier | Hot tier |
|---|---|---|
| Stream | Log segments | Tail cache of recent records |
| Collection (vectors) | Lance with an IVF index | HNSW graphs on NVMe, derived from Qdrant's design |
| Collection (text) | Tantivy splits | Pinned splits and an in-memory tail index |
| Table | Iceberg | Sorted projections, in the style of ClickHouse |
| Graph | CSR and CSC sidecars | In-memory adjacency |
Query nodes are chosen by namespace and object affinity, so the same node keeps serving the same hot data. When a node dies, its objects move to another node, which warms from the bucket or from prebuilt artifacts.
Why the metastore is Raft
Streams generate high-rate metadata: offset assignment on every flush, consumer offset commits, leases. A conditional PUT to object storage takes tens to hundreds of milliseconds and contends per key, which cannot sustain that. So metadata lives in a small embedded Raft group (openraft), snapshotted to the bucket. Object data never flows through it.
Read more in §03 Storage formats and §04 Hot tier and caching.