Storage and the hot tier

What lives in the bucket, and what compute nodes keep to make it fast.

Every object in Operon has two tiers. The durable tier is an open format in object storage and the only source of truth. The hot tier is derived, node-local and rebuildable. Losing the hot tier changes latency, never correctness.

The bucket

Everything at rest lives under one prefix in one bucket:

s3://<bucket>/<cluster_prefix>/
  meta/snapshots/…                          # metastore snapshots
  wal/<class>/<node_id>/<ulid>.wal          # WAL objects, many partitions and namespaces each
  ns/<namespace_id>/
    streams/<stream_id>/<partition>/<base_offset>-<ulid>.seg
    collections/<collection_id>/
      lance/…                               # Lance dataset
      text/splits/<ulid>.split              # Tantivy splits
      manifests/<version>.pb                # immutable collection manifests
      hot/…                                 # optional prebuilt hot-tier artifacts (HNSW)
    graphs/<graph_id>/
      idmap/…  adj/…  manifests/…
    pk/<object_id>/…                        # primary key → row location
    durable/wf/<origin>                     # one document per workflow origin
  warehouse/<namespace_id>/<table_id>/      # Iceberg tables, catalogued in Lakekeeper

All data objects are immutable and named by ULID or version. Only the metastore's pointers and the Iceberg catalog's pointers move. Operon relies on conditional writes (If-None-Match, If-Match), which S3, GCS and Azure all support.

The hot tier

ObjectDurable tierHot tier
StreamLog segmentsTail cache of recent records
Collection (vectors)Lance with an IVF indexHNSW graphs on NVMe, derived from Qdrant's design
Collection (text)Tantivy splitsPinned splits and an in-memory tail index
TableIcebergSorted projections, in the style of ClickHouse
GraphCSR and CSC sidecarsIn-memory adjacency

Query nodes are chosen by namespace and object affinity, so the same node keeps serving the same hot data. When a node dies, its objects move to another node, which warms from the bucket or from prebuilt artifacts.

Why the metastore is Raft

Streams generate high-rate metadata: offset assignment on every flush, consumer offset commits, leases. A conditional PUT to object storage takes tens to hundreds of milliseconds and contends per key, which cannot sustain that. So metadata lives in a small embedded Raft group (openraft), snapshotted to the bucket. Object data never flows through it.

Read more in §03 Storage formats and §04 Hot tier and caching.

On this page