The log and consistency

How a write travels through Operon and why any read can see it.

Every write in Operon, whether a native write, a Flight DoPut, an Elasticsearch _bulk or a Qdrant upsert, lands in a stream first. Everything else is derived from that log.

The write path

A gateway translates the request

Whatever protocol the client speaks, the gateway turns it into a logical write against a stream, explicit or implicit.

A log node appends it to the WAL

Any log node can accept writes for any partition. It buffers batches from many partitions, writes one WAL object to the bucket with a create-only conditional PUT, and asks the metastore to assign offsets.

The client gets a consistency token

The acknowledgement carries the dense offsets the metastore assigned, as a token: a set of (stream, partition, offset).

In the background, workers build the derived forms (Lance fragments, Tantivy splits, Iceberg data files, adjacency sidecars) and commit each one together with the offset it reflects.

WAL classes

Each stream chooses how its writes become durable.

ClassHowProduce p99 (target)Best for
standardWAL objects on regional object storage400–600 msBulk ingest, collection and table ingest, cost-first topics
expressWAL objects written to three zonal buckets, acknowledged on two20–50 msLatency-sensitive topics without stateful disks
quorumA three-node Raft journal on local NVMe, offloaded to the bucket3–10 msThe lowest latency, on-premises, clouds without zonal object storage

All three classes survive the loss of any node and any single availability zone. Today only standard is built; express arrives in M3 and quorum in M5.

Latency figures are design targets

They come from reference systems and published numbers, not from Operon measurements.

Consistency tokens

A consistency token names the offsets a reader needs to see. Every write returns one:

{ "base_offset": 41, "last_offset": 42, "token": [{ "stream": 1, "partition": 0, "offset": 42 }] }

Pass the token to a read on any object derived from that stream and the read is guaranteed to reflect it.

Reads merge the tail

Indexes lag the log by the time it takes a worker to apply a batch. Operon closes that gap on every read:

read = durable or hot state at the applied offset  ∪  tail (applied offset, requested offset]

The tail is the part of the log that is committed but not yet in the durable indexed form. Query nodes keep it in memory and merge it into results, so the default read is strong: it sees every write acknowledged before it began. An eventual read skips the tail for lower latency.

ScopeGuarantee
Within a stream partitionTotal order; acknowledged writes are durable per WAL class
Single-object reads (default)Strong, through the tail merge
Cross-object readsA snapshot per object; with a token, every object derived from those streams reflects its offsets
Multi-record writesAtomic per request within one stream
External Iceberg readersIceberg snapshots at commit cadence, without the tail
Not providedMulti-object serializable transactions; interactive OLTP transactions

Read more in §01 Architecture and §02 Stream engine.

On this page