Loam is pre-alpha: the engine core runs today; Live and Durable are in progress. See the roadmap

Blog/testing

Crash gates: how we test kill -9 at every step

Named failpoints, a store fault matrix and a seeded linearizability simulation. This is what "no acknowledged write is ever lost" means in Loam's test suite.

Crash gates: how we test kill -9 at every step
On this page
  1. Gate 1: kill -9 at every named step
  2. Gate 2: every component against every store fault
  3. Gate 3: linearizability, checked
  4. Why it's worth the machinery

A storage engine makes one promise above all: once it acknowledges a write, the write is not lost, and no reader ever sees half of a commit. Loam's write path touches object storage, a Raft metastore, a primary-key index and two file formats, so "the process died here" has many possible values of "here".

We don't rely on hoping to hit them. Every milestone closes on exit gates, and the foundation milestone's three gates run on every change to the write path. This post walks through them.

Gate 1: kill -9 at every named step

Every step between "bytes written" and "commit visible" has a named failpoint (the fail crate, compiled in only with the failpoints feature, so release builds carry none). The crash test:

  1. starts operon dev as a child process with one failpoint armed to call std::process::abort() on its n-th hit;
  2. has a client produce numbered records to several partitions while a link materializes them;
  3. lets the process die at that point, then restarts it without failpoints and lets background work settle;
  4. checks that every acknowledged record is readable exactly once at its acknowledged offset, that offsets are dense, that the link's output equals the log exactly once, and that the metastore's invariants hold.

The foundation's failpoints cover the WAL (wal.after_put, wal.after_commit), the segmenter (seg.after_put, seg.after_swap), link apply (link.after_data_put, link.after_manifest_put, link.after_cas), GC (gc.after_delete), metastore snapshots (meta.snapshot.after_put, meta.snapshot.after_pointer) and retention (retention.after_trim).

Collection storage added a collection scenario with its own failpoints: collection.after_lance_commit, collection.after_split_put, collection.after_manifest_put, collection.after_cas, collection.after_pk_write, plus the two index-build points. After each abort, the committed collection must equal the fold of its log: every acknowledged insert, upsert, patch and delete, applied in order, and nothing else.

Named points only cover the places someone thought of. So the gate also runs a random SIGKILL loop under load, killing the process at arbitrary moments (20 kills by default) and running the same checks after each restart.

When the foundation milestone closed, three full runs passed 36 of 36 tests: 33 failpoint aborts and 60 random SIGKILLs, with no lost or duplicated record (foundation exit report).

Gate 2: every component against every store fault

Object storage fails in ways a local disk doesn't. A PUT can return an error after it has actually been applied. A conditional write can hit a 412. A request can stall for seconds. FaultyStore injects these faults deterministically, and the fault-matrix test crosses:

  • components: writer flush, reader fetch, segmenter swap, retention trim, link commit, GC pass, metastore snapshot, and, since collection storage, the collection commit;
  • store operations: Put, PutCreate, PutIfMatch, Get, Delete, List;
  • faults: Error, ErrorAfterApply, Precondition (409 on create-only, 412 on compare-and-swap), and a 2-second Delay;

each injected on the first call and again on the second.

Every cell's outcome must match a committed expectation table: retried, deferred, surfaced as a retryable error, or no effect. An error that never reached the fault fails the gate. After every cell, the test checks again that no acknowledged record was lost or changed and no failed write is visible. When the foundation closed, the matrix had 336 cells, all asserted.

Gate 3: linearizability, checked

The two places where Loam needs linearizability are the per-partition offset sequencer and the manifest-pointer compare-and-swap. operon-sim checks both. It runs a three-node metastore over an in-process router, together with writers, a reader and a worker doing segmenting, retention, link apply and GC. It uses seeded workloads, random store faults, node isolation and healing, and worker crashes. It records every operation's history and checks it with its own Wing–Gong–Lowe linearizability checker.

The checker is tested against hand-written histories too: a stale read, a lost update, a duplicate offset, a gap, and indeterminate operations taken either way. CI sweeps 32 seeds of 300 steps on every run. The foundation exit ran an extra 64-seed sweep: 5,849 acknowledged appends and 79 indeterminate operations, with zero violations.

Why it's worth the machinery

These gates catch the bugs that code review misses. A GC that deletes a young WAL object. A pointer that dangles after a crash between two PUTs. A retried write committed twice. They also make the design honest: a rule like "the manifest CAS is the only commit point" (One manifest, two formats) is only as good as the test that kills the process on either side of it.

The gates are plain cargo test targets in the engine repo. Run them yourself:

cargo test -p operon --features failpoints --test crash
cargo test -p operon --test fault_matrix
SIM_SEEDS=64 cargo test -p operon-sim --release --test sim

More from the blog