§28 Loams Postgres: a Neon Fork with Loams's Control Plane and Loams's WAL

CloudNativePG (Postgres 17, Barman Cloud PITR to RustFS) for the showcase apps; `dina-kar/neon` with Loams as the primary control plane (compute spec, storage-controller hooks, proxy auth API); PgDog (AGPL-3.0) as an unmodified router; Loams's WAL replacing safekeepers: the safekeeper protocol in `loams-safekeeper`, TiKV 1PC as the quorum hot tier, the Loams log as the bucket copy; the pgbench merge gate; the fork-maintenance plan; phases P1–P5 (D230–D241)

Status: Approved direction · 2026-09-29. On 2026-09-29 the owner amended §23 (PR #111) with these decisions:

  • The showcase apps run on CloudNativePG.
  • Neon is forked, with Loams's control plane as the primary control plane.
  • PgDog does the routing, and only as an unmodified separate service.
  • Loams's own WAL replaces Neon's safekeepers.

The owner then chose the WAL design:

  • Option A. TiKV is the quorum hot tier, and Loams's log group-commits to the bucket behind it.
  • A Loams crate speaks Neon's safekeeper protocol, so walproposer and the pageserver stay unmodified.
  • A pgbench merge gate gates the switch-over.

These are decisions D230–D236, all owner-approved on 2026-09-29. The WAL replacement (D233) is approved direction: it becomes the default only if the benchmark gate passes. The concrete WAL design is D237–D241. They are Loams's proposals within the owner's decisions and are marked as such. Open questions are Q110–Q119. Arm A of the WAL (§7.2, 2026-09-30) adds D263–D269 and D271 (D270 is the CloudEvents decision; D263, compio for the data path, is the owner's) and Q261–Q264.

This document supersedes §23's D149 (Neon for the showcase apps) and D151 (fork only when needed). It amends D150 (the control plane grows from a client into the primary control plane) and D153 (PgDog, not Loams's pg listener, splices Postgres connections to computes). D148, D152, D154 and D155 stand. D156 and D157's WeSQL row are not affected.

Markers:

  • (spike) means measured on 2026-09-29 on the development machine. That machine is a 14-core, 15 GB laptop with consumer NVMe and btrfs, and other sessions' builds were running, with a load average between 1.5 and 24. The notes are in .superpowers/research/loams-postgres-spike.md, local and not committed. Spike numbers compare paths on the same hardware. They are not production figures.
  • (verify) means not checked against a primary source.
  • (estimate) means computed, not measured.
  • Paths of the form neon/… point into dina-kar/neon at fa504217c (2026-08-31, identical to upstream main).

1. Summary

#DecisionStatus
D230The showcase apps run on CloudNativePG (Apache-2.0, v1.30.1) with plain Postgres 17 (17.11). Backups and PITR go to RustFS through the Barman Cloud CNPG-I plugin (v0.15.0), not the in-tree barmanObjectStore, which is deprecated and removed in 1.31. Replaces D149Approved (owner, 2026-09-29)
D231Loams Postgres is a fork of Neon, dina-kar/neon (created 2026-09-29 as a GitHub fork of neondatabase/neon, Apache-2.0). Loams owns its releases, its Postgres patch rebases and its images. Nothing is posted upstream. Replaces D151Approved (owner, 2026-09-29)
D232Loams's control plane is Loams Postgres' primary control plane. It replaces Neon's closed one: tenant, timeline and branch lifecycle; compute specs and compute start; the storage controller's hooks; the proxy's auth API when the proxy is used (§5). Amends D150Approved (owner, 2026-09-29)
D233Loams's WAL replaces Neon's safekeepers, behind a feature, and becomes the default only when the pgbench gate passes: p99 commit latency no worse than the safekeeper baseline, and throughput no worse (§6, §7)Approved direction (owner, 2026-09-29), gated on benchmarks
D234Option A with TiKV as the quorum hot tier. A Postgres WAL record is acknowledged once it is durable in TiKV (Raft quorum). Loams's log then group-commits it to the bucket. TiKV's copy is trimmed after the bucket upload and the pageserver's remote_consistent_lsn pass it. The target is single-digit-ms commit acksApproved (owner, 2026-09-29)
D235Loams speaks Neon's safekeeper protocol through a clean Rust API. A Loams crate, loams-safekeeper, implements the proposer–acceptor protocol (v3), the replication protocol the pageserver and computes read from, and the storage-broker publication. Walproposer and the pageserver run unmodified. P4 splits into P4a (the protocol crate and its TiKV backend), P4b (the benchmark harness) and P4c (the switch-over behind a feature, merged only when the gate passes)Approved (owner, 2026-09-29)
D236PgDog routes, as an unmodified separate service only (AGPL-3.0, v0.1.60). It is never linked, vendored or patched. Loams generates its config. The fallbacks are Neon's proxy (Apache-2.0) and PgBouncer (ISC). Amends D153 for Postgres OLTPApproved (owner, 2026-09-29)
D237TxnKV with 1PC and async commit, not RawKV, for the WAL hot tier. Only a transaction can check the term fence and write the WAL in one atomic step. RawKV needs a separate CAS round trip, and it was not faster in the spike (§6.4)Proposed (Loams's design within D234)
D238Key layout: a dedicated keyspace loams_pgwal; per timeline, a head key followed by chunk keys by LSN; one split boundary per timeline; batches of 128 KiB chunks up to 1 MiB per transaction (§6.5)Proposed
D239One logical acceptor per timeline, served by a stateless, shared WAL service (the wal role, one pool per AZ). It is not a compute sidecar. The compute's neon.safekeepers names one endpoint (quorum of one). Durability comes from TiKV's Raft quorum, not from three acceptors (§6.3)Proposed
D240The pageserver reads WAL from the WAL service over the unchanged START_REPLICATION interpreted protocol. The service serves from its tail cache, then TiKV, then the Loams log in the bucket. Trimming waits for both the bucket commit and the pageserver's remote_consistent_lsn (§6.7)Proposed
D241The fork's Postgres is rebased every quarter. Postgres 17 and 18 are supported, and 16 until it reaches end of life. The minor-release merge runs in CI on upstream's release day, and the compute images are published within a week (§10)Proposed

2. Goals and non-goals

2.1 Goals

  1. Real Postgres for the showcase apps now. Plain Postgres on CloudNativePG, with PITR to the same RustFS store. It takes no fork work (D230).
  2. Serverless, branchable Postgres on the bucket as a Loams product. Neon's storage (pageserver, layers in the bucket, copy-on-write branches) runs under Loams's control plane, with no dependency on a vendor's closed services (D231, D232).
  3. A WAL service that beats safekeepers where it counts.
    • Operations: no stateful safekeeper fleet, and no membership migrations.
    • Durability: the bucket copy is seconds old, not segments old.
    • Cost: TiKV nodes are shared with the metastore and Loams Live.
    • Latency: commit latency stays at or under the baseline, and the gate enforces it (D233).
  4. License-clean. Apache-2.0 for everything Loams links (D11). PgDog, which is AGPL-3.0, runs only as an unmodified process (D236).

2.2 Non-goals

  • Changing walproposer or the pageserver's WAL receiver. Option C (a Loams protocol inside walproposer) is rejected (§6.2).
  • Merging branches. As in §23: Postgres cannot merge timelines (D155).
  • Neon's hosted extras. The console, the Data API, Neon Auth and autoscaling VMs are out of scope. Scale-to-zero is Q113.
  • Upstream contributions. Nothing is posted to neondatabase, pgdogdev or cloudnative-pg without the owner's approval.
  • Modifying TiKV or PD. D126 holds: official releases, configured, never patched. Client-side work goes into Loams's client-rust fork.

3. Architecture

  apps / agents / psql                                   showcase apps (Plane, Zulip, GlitchTip,
        │ postgres://…/<db>[__<branch>]                    Keycloak, OpenFGA)
        ▼                                                      │ postgres://
 ┌──────────────────────────────┐                     ┌────────┴──────────────────────────────┐
 │ PgDog (AGPL-3.0, unmodified, │                     │ CloudNativePG cluster (P1, D230)      │
 │ separate pod): pools, routes │                     │ Postgres 17.11, primary + replica     │
 │ by database name, balances   │                     │ Barman Cloud plugin ──► RustFS (PITR) │
 │ reads across replicas        │                     └───────────────────────────────────────┘
 └──────┬───────────────────────┘
        │ config + RELOAD from Loams's control plane
        ▼
 ┌──────────────────────────────┐   spec (GET …/computes/{id}/spec)   ┌─────────────────────────────┐
 │ Loams Postgres compute       │◄────────────────────────────────────┤ Loams control plane (D232)  │
 │ (fork image, compute_ctl,    │                                     │ in `loams`, state in TiKV   │
 │ walproposer, unmodified)     │                                     │ tenants / timelines /       │
 └──────┬──────────────┬────────┘                                     │ branches / computes;        │
        │ safekeeper   │ GetPage@LSN                                  │ /notify-attach,             │
        │ protocol v3  │                                              │ /notify-safekeepers,        │
        ▼              ▼                                              │ proxy auth API (optional)   │
 ┌──────────────────────────┐   ┌─────────────────────────────────┐   └──────┬──────────────────────┘
 │ Loams WAL service (P4)   │   │ pageserver (fork, unmodified    │          │ mgmt API
 │ `loams-safekeeper`,      ├──►│ WAL receiver)                   │◄─────────┤ storage controller
 │ role `wal`, stateless,   │   │ layers ──► bucket               │          │ (fork; its DB on CNPG,
 │ one pool per AZ          │   └─────────────────────────────────┘          │ TiKV later, P5)
 │ ├─ 1PC txn ──► TiKV keyspace `loams_pgwal` (3 replicas, leader in the compute's AZ)
 │ └─ group commit (250 ms) ──► Loams log (loams-log `standard`) ──► bucket
 │ publishes SafekeeperTimelineInfo ──► storage broker (fork, unmodified)
 └──────────────────────────┘
      stock safekeepers ×3 remain the default until the P4c gate passes (P2)

 RustFS: bucket `loams` (Loams, including the WAL log), bucket `neon` (layers), bucket `cnpg` (Barman)

3.1 Components

ComponentSourceLicenseRole
CloudNativePG operatorcloudnative-pg/cloudnative-pg v1.30.1 (2026-09-23)Apache-2.0The showcase's Postgres (P1) and the storage controller's database (P2)
Barman Cloud plugincloudnative-pg/plugin-barman-cloud v0.15.0 (2026-09-03)Apache-2.0WAL archiving and base backups to RustFS; PITR
Loams Postgresdina-kar/neon (fork of neondatabase/neon at fa504217c)Apache-2.0Pageserver, storage broker, storage controller, compute_ctl, compute images, and stock safekeepers until P4c
Loams Postgres' Postgresdina-kar/postgres (to fork from neondatabase/postgres in P2, Q110)PostgreSQL LicenseThe vendor/postgres-v1x submodules
Loams control planeloams, extending §23's loams-neonApache-2.0§5
Loams WAL serviceloams-safekeeper (new, P4a)Apache-2.0§6
PgDogpgdogdev/pgdog v0.1.60 (2026-09-24), image pinned by digestAGPL-3.0§8. Unmodified, separate process
TiKV and PDv8.5.x, official releases (D126)Apache-2.0WAL hot tier (a keyspace of the existing cluster)
RustFS1.0.0Apache-2.0The object store (D61)

4. CloudNativePG for the showcase apps (D230, P1)

Why. The showcase apps need Postgres now. Neon's fork work (§10) and the WAL service (§6) take quarters, and D-SC-12's Postgres front end over TiKV takes years. CloudNativePG is plain Postgres with a mature operator, PITR to object storage, and no fork. That is the same "plain Postgres as the exit" that D151 described.

Verified 2026-09-29:

FactEvidence
Apache-2.0; CNCF Sandbox since 2025-01-21 (incubation applied, not yet granted)gh api repos/cloudnative-pg/cloudnative-pg (license.spdx_id); cncf.io/projects/cloudnativepg
Latest v1.30.1 and v1.29.3 (both 2026-09-23). 1.30.x supports PG 14–18 and Kubernetes 1.34–1.36, EOL about Dec 2026. 1.29.x reaches EOL 2026-09-29GitHub releases; cloudnative-pg.io/docs/devel/supported_releases
Default operand 18.6-system-trixie. The 17.11-{minimal,standard,system}-trixie images exist. system images are deprecatedv1.30.1 notes; ghcr manifests; cloudnative-pg/postgres-containers
In-tree barmanObjectStore deprecated since 1.26, removal planned for 1.31.0v1.30.1 release notes; docs/backup
Barman Cloud plugin: Apache-2.0, v0.15.0 (Barman 3.20.0). ObjectStore CRD barmancloud.cnpg.io/v1 with endpointURL, endpointCA, s3Credentialsgh api repos/cloudnative-pg/plugin-barman-cloud/releases; plugin docs
S3-compatible stores: MinIO is documented. Newer boto3 checksums can fail with x-amz-content-sha256 errors, and the fix is AWS_REQUEST_CHECKSUM_CALCULATION=when_required and AWS_RESPONSE_CHECKSUM_VALIDATION=when_required. No RustFS report found, so P1 tests it (§23 §9.1 hit the same header with RustFS)plugin docs, object_stores

Layout.

  • One Cluster, commons-pg, for the suite. It has instances 2 (a primary and a streaming replica), image ghcr.io/cloudnative-pg/postgresql:17.11-standard-trixie pinned by digest, and one database and owner role per app. This keeps §22 §6.4's "one Postgres 17 with a database per app".
  • Zulip needs PGroonga, which the standard image lacks. It gets an image built on the standard image with PGroonga, or a separate Cluster (Q111).
  • Backups:
    • An ObjectStore named commons-rustfs sets destinationPath: s3://cnpg/commons-pg, endpointURL: http://rustfs:9000, credentials from a Secret, wal.compression: zstd and the two checksum variables above.
    • The Cluster declares plugins: [{name: barman-cloud.cloudnative-pg.io, isWALArchiver: true, parameters: {barmanObjectName: commons-rustfs}}].
    • A ScheduledBackup with method: plugin runs daily.
  • PITR is a new Cluster with bootstrap.recovery.source: origin and recoveryTarget.targetTime, and externalClusters: [{name: origin, plugin: {name: barman-cloud.cloudnative-pg.io, parameters: {barmanObjectName: commons-rustfs, serverName: commons-pg}}}]. P1's test restores to a timestamp between two known writes.
  • Dev without Kubernetes is not supported by CNPG. The suite's compose keeps a plain postgres:17.11 container. P1 adds deploy/cnpg/ with a kind or k3d manifest set for the PITR test.

Loams integration. The same as §23 for everything that addresses "a Postgres backend": the x/ record carries engine = postgres (D152), and the pgoutput bridge (D154) runs against CNPG's primary, with wal_level = logical set in the Cluster's postgresql.parameters. Branch-per-workspace (D155) needs Loams Postgres and is not offered on CNPG.

5. The Loams control plane (D232)

§23's loams-neon was a client of Neon's APIs. Here Loams is the control plane that Neon's components call, so it serves as well as calls. The endpoints below come from the fork's code, which is the only specification.

5.1 What Loams calls

TargetCallPurpose
Storage controller (or a pageserver directly in dev)POST /v1/tenant/{t}/timeline (create or branch: new_timeline_id, ancestor_timeline_id, ancestor_start_lsn, pg_version), DELETE …/timeline/{tl}, PUT /v1/tenant/{t}/location_config, GET /v1/tenant/{t}/timelineTenant, timeline and branch lifecycle (§23 §6.1)
Storage controllerPOST/GET /control/v1/safekeeper[/:id], …/scheduling_policy, …/safekeeper_migrateStock safekeepers only, from P2 until P4c. The Loams WAL service needs none of these (§6.3)
compute_ctl (per compute, JWT from compute_ctl_config's JWKS)POST /configure (ConfigurationRequest{spec, compute_ctl_config}), GET /status, POST /terminate, POST /promotePush a new spec (for example a new pageserver after a migration), stop, promote a replica
ComputeRuntime (§23 §6.1)start, stop, statusStart a compute pod with --control-plane-uri pointing at Loams
PgDog admin databaseRELOAD after writing pgdog.toml and users.tomlRouting (§8)

5.2 What Loams serves

CallerEndpointRequest → responseLoams's handler
compute_ctl at startGET {cp}/compute/api/v2/computes/{compute_id}/spec, Authorization: Bearer $NEON_CONTROL_PLANE_TOKEN. It retries up to 3 times on 502/503→ ControlPlaneConfigResponse{spec?: ComputeSpec, status: "empty" | "attached", compute_ctl_config} (neon/libs/compute_api/src/responses.rs)Builds ComputeSpec from the x/ and X/ records: tenant_id, timeline_id, mode, pageserver_connection_info, safekeeper_connstrings (+ safekeepers_generation), storage_auth_token, roles, databases, settings (wal_level = logical, synchronous_standby_names = 'walproposer')
Storage controllerPUT {cp}/notify-attach with {tenant_id, preferred_az?, stripe_size?, shards: [{node_id, shard_number}]}2xx; 423 Busy; 429 SlowDownRecords the pageserver per shard and pushes /configure to the tenant's computes
Storage controllerPUT {cp}/notify-safekeepers with {tenant_id, timeline_id, generation, safekeepers: [{id, hostname?}]}as aboveStock safekeepers only (P2–P4c): pushes a new safekeeper_connstrings and generation to computes
Neon proxy (only if Q113 chooses it)GET {cp}/get_endpoint_access_control?session_id&application_name&endpointish&role → {role_secret, allowed_ips?, allowed_vpc_endpoint_ids?, block_public_connections?, rate_limits, …}; GET {cp}/wake_compute?… → {address, server_name?, aux: {endpoint_id, project_id, branch_id, compute_id, cold_start_info}}; GET {cp}/endpoints/{endpoint}/jwksneon/proxy/src/control_plane/client/cplane_proxy_v1.rs, messages.rsRole secrets come from the credential store (Q30). wake_compute starts a compute through ComputeRuntime and waits for /status

The pageserver's upcalls (POST /upcall/v1/re-attach, POST /upcall/v1/validate) go to the storage controller, which issues the generations that fence S3 writes (Neon RFC 025). Loams does not replace them in P2. The storage controller runs from the fork, with its Postgres database on CNPG (Q48 answered: a small CNPG Cluster). A TiKV-backed replacement is P5 (§9).

5.3 State

RecordKey (loams-meta-tikv)Change from §23
Databasex/<ns>/<db>Adds engine = loams-pg | postgres, wal = safekeepers | loams (per database, so P4c can switch one database at a time), storage_controller endpoint
BranchX/<ns>/<db>/<branch>Adds compute_id, pageserver shard map from notify-attach
ComputeC/<compute_id> (new; prefix checked as unused in crates/loams-meta-tikv/src/keys.rs at e69ab38)Spec version, status, endpoint, JWT key id, created and last-seen times

Writes are compare-and-set on a version, as in §23 §6.2. The spec endpoint is idempotent: it returns the same spec for the same compute id until the branch record changes.

6. Replacing the safekeepers (D233–D235, D237–D241)

6.1 What safekeepers do today

Read from the fork (neon/pgxn/neon/walproposer*.c, neon/safekeeper/src/, neon/pageserver/src/tenant/timeline/walreceiver/, neon/docs/walservice.md, RFCs 013, 035 and 041):

  • Commit path.
    • walproposer is a background worker that poses as a synchronous standby (synchronous_standby_names = 'walproposer').
    • A committing backend waits in stock SyncRepWaitForLSN.
    • walproposer computes commitLsn as the quorum (n/2 + 1) of the safekeepers' AppendResponse.flush_lsn, counting only WAL of its own term, and releases the waiters through ProcessStandbyReply.
    • A commit returns when a majority has fdatasynced its WAL (walproposer.c GetAcknowledgedByQuorumWALPosition; walproposer_pg.c walprop_pg_process_safekeeper_feedback).
  • Safekeeper write path.
    • Each AppendRequest (at most 128 KiB) is written without fsync.
    • One fdatasync covers everything queued when the inbound queue drains, or after at most 1 s (receive_wal.rs FLUSH_INTERVAL).
    • Segments are preallocated, so fdatasync suffices.
    • The control file is saved every 300 s or every segment, and at once on votes and term changes.
  • Offload.
    • One elected offloader per timeline uploads full 16 MiB segments once commit_lsn passes them.
    • The current partial segment is uploaded within 15 minutes (wal_backup.rs, wal_backup_partial.rs).
    • Local WAL is removed below min(remote_consistent_lsn, backup_lsn, commit_lsn, flush_lsn).
  • Consumers.
    • The pageserver finds safekeepers through the storage broker (SafekeeperTimelineInfo, pushed every 1 s).
    • It picks the one with the highest commit_lsn, preferring its own AZ.
    • It reads with START_REPLICATION PHYSICAL using the interpreted protocol (protobuf and zstd; records decoded and sharded on the safekeeper; libs/wal_decoder).
    • It reports remote_consistent_lsn back in a 'z' status update.
  • The machinery around it:
    • term histories and truncation;
    • peer recovery;
    • timeline eviction;
    • partial-segment GC;
    • two-phase membership changes with generations and pull_timeline;
    • a storage-controller scheduler for safekeepers.

What the repository does not say. Neither the fork nor Neon's site publishes a safekeeper commit latency. Neon's architecture page says only that "commit latency is primarily quorum/network-bound" (neon.com/docs/introduction/architecture-overview). The one third-party figure found is ClickHouse's PostgresBench (2026-07-21): Neon at 16 vCPU ran 7,802 TPS with p50 32.8 ms and p99 56.3 ms. That is pgbench latency under load, not a commit round trip (clickhouse.com/blog/postgresbench-ha). The production baseline must therefore be measured (P4b). The estimate for three safekeepers across AZs is one cross-AZ round trip plus an NVMe fdatasync, about 1–3 ms at p50 (estimate, §6.6).

Measured on this machine (spike). The stack was deploy/neon/ with one safekeeper (fsync on, WAL in the container filesystem), Postgres 16.9, and pgbench run inside the compute:

WorkloadClientsTPSp50p90p99
Single-row INSERT, one commit per transaction12203.57 ms5.79 ms22.4 ms
Same161 4748.53 ms18.3 ms40.3 ms
Same, synchronous_commit = off16 7430.09 ms0.32 ms0.81 ms
TPC-B (-s 10)11117.98 ms11.0 ms23.2 ms
TPC-B (-s 10)1649524.5 ms59.2 ms133 ms

So the safekeeper ack costs about 3.5 ms at p50 on this hardware, almost all of it the fdatasync of a consumer SSD under btrfs.

6.2 Options

OptionWhat changesForAgainst
A. Loams's hot tier acks at quorum, and Loams's log group-commits to the bucketA WAL store under the safekeeper roleKeeps a low-latency ack; reuses TiKV and the Loams logNeeds a protocol front end: B or C
B. The safekeeper wire protocol, implemented by LoamsNew Loams service; walproposer, pageserver and broker unchangedNo change to the patched compute; upgrades are drop-in; the benchmark compares like with likeLoams must match a subtle protocol (terms, truncation, interpreted sending)
C. walproposer speaks Loams's protocolThe C extension in pgxn/neon; the pageserver's receiverCould drop the Paxos layer completelyForks the most delicate C code in Neon; every Postgres rebase carries it; the pageserver changes too
Pure bucket WAL (standard or express class)Ack after a bucket PUT and a metastore commitCheapest; no hot tierMeasured 22 ms p50 and 148 ms p99 at a 1 ms flush interval (§6.4). S3 Standard is tens of ms per PUT (AWS: "median latencies are often in the tens of milliseconds"). express targets 20–50 ms (§02 §2). All of these are an order of magnitude over safekeepers for OLTP

Chosen: A behind B (D234, D235, owner-approved). TiKV provides the quorum, and a Loams crate speaking the safekeeper protocol sits in front. C is rejected: it moves the risk into C code that every rebase must carry.

6.3 One logical acceptor per timeline (D239)

Walproposer is Paxos over n independent acceptors. Three Loams acceptors writing to one TiKV store would triple every write for no gain, because TiKV already replicates across three AZs. So:

  • The compute's neon.safekeepers names one endpoint. It is a per-AZ Service address of the WAL pool, so walproposer runs with n = 1 and quorum 1. Neon runs this way in development today: deploy/neon has one safekeeper, and so does the spike.

  • The acceptor's state lives in TiKV, not in a process: term, term history, flush_lsn, commit_lsn, peer_horizon_lsn, backup_lsn and remote_consistent_lsn. Any instance of the pool can serve any timeline. An instance that picks up a connection behaves exactly like a safekeeper that restarted with its disk intact: it answers the greeting from TiKV's state, refuses a vote for a term it has already voted in, and accepts ProposerElected from the current proposer.

  • Losing an instance is a reconnect, not a failover. walproposer reconnects after neon.safekeeper_reconnect_timeout (default 1 000 ms; a PGC_SIGHUP GUC with minimum 0, walproposer_pg.c; Loams sets 100 ms). Nothing is copied, and nothing needs electing on the Loams side.

  • What disappears:

    • membership changes and generations (the generation is fixed at 1 with one member, the Service);
    • pull_timeline and peer recovery;
    • timeline eviction and partial-segment uploads;
    • the storage controller's safekeeper scheduler.

    That is thousands of lines of Neon's safekeeper and controller code, and the operational runbooks that go with them.

  • Sidecar rejected. A WAL process in the compute pod saves one in-AZ hop (about 0.1–0.2 ms, estimate). But it dies with the compute, while the pageserver still has to read the tail after the compute stops. It also multiplies instances by computes. The shared pool is stateless and needs no compute integration.

6.4 The TiKV API (D237)

The fence. A WAL write is valid only if the timeline's term is still the proposer's term when it lands. This is check_fence's rule in loams-meta-tikv (leases.rs): read the fence key for update in the same transaction as the write, compare the epoch, and abort on mismatch. The WAL service applies it to the timeline's head key:

append(tl, term, begin_lsn, bytes, commit_lsn):
  txn (optimistic, use_async_commit, try_one_pc):
    head = get(H(tl))                      # fence read
    if head.term != term  -> abort: reply AppendResponse{term: head.term}   # deposed
    if end_lsn <= head.flush_lsn -> fully covered retry: write no chunk; put the head only if
                                    commit_lsn grew; reply with the existing flush_lsn
    if head.flush_lsn > begin_lsn -> drop the overlapping prefix (same term: same bytes)
    if head.flush_lsn < begin_lsn -> error (gap), the proposer reconnects
    put(W(tl, max(begin_lsn, head.flush_lsn)), remaining bytes)   # one key per ≤128 KiB chunk
    put(H(tl), head{flush_lsn = end_lsn, commit_lsn = max(head.commit_lsn, commit_lsn)})
                                           # end_lsn > head.flush_lsn here: flush_lsn only advances
  commit (1PC when H and W are in one region, else async commit)
  -> AppendResponse{term, flush_lsn = end_lsn, commit_lsn, …}
  • The head is written, not only read. A concurrent term bump (a vote handled by another instance) therefore conflicts write-write in either order:

    • If the bump committed before this transaction's start_ts, the fence read sees the new term.
    • If it committed after, one of the two prewrites fails.

    No timing assumption is involved, which a lease-and-wait scheme would need.

  • A vote is the same transaction shape: read the head, grant the vote only if head.term < term, and write the head with the new term. The term is durable before VoteResponse, as in handle_vote_request.

  • ProposerElected truncates at start_streaming_at, in one transaction. A chunk key covers a range of LSNs, so the chunk that contains start_streaming_at is rewritten with only its bytes below that LSN; every chunk starting at or above it is deleted; and the head gets the new term history and flush_lsn = start_streaming_at. Committed WAL is never truncated (start_streaming_at ≥ commit_lsn is asserted, as in handle_elected).

Why not RawKV. A RawKV write cannot check another key. The fenced raw scheme is "put the chunk, then CAS the head", which is two sequential Raft writes in atomic-CAS mode. A raw put without the CAS lets a deposed proposer's write be acknowledged. With one acceptor, a deposed compute would then count a commit that the new term truncates: an acknowledged transaction lost. The spike measured the options on one TiKV store (playground v8.5.8, API v2, default configuration, 1 000 sequential operations, two runs):

8 KiB WAL chunkp50 (run 1 / run 2)p99 (run 1 / run 2)
TxnKV 1PC + async commit: get head, put chunk (the design)6.55 / 7.14 ms25.2 / 26.0 ms
TxnKV 1PC blind put (no fence; the floor)4.62 / 10.9 ms18.3 / 46.5 ms
TxnKV 2PC: get head, put chunk25.8 / 19.0 ms57.4 / 74.2 ms
RawKV put (no fence)14.9 / 6.44 ms50.5 / 33.3 ms
RawKV put, atomic-CAS mode (no fence)10.9 / 13.8 ms55.3 / 59.8 ms
RawKV CAS(head) + put (fenced)29.8 / 13.8 ms84.9 / 595 ms
128 KiB chunk, TxnKV 1PC: get head, put chunk7.17 / 7.00 ms35.9 / 21.1 ms

On this machine RawKV was not faster than a 1PC transaction. The fenced raw scheme was 2–4× slower at p50, and 2PC was about 3× slower than 1PC. The single-store playground has no network replication, so every figure here is "local Raft fsync plus apply plus client". The absolute values are inflated by the shared, loaded machine and must be re-measured on a three-node cluster (Q114). The ordering is the result: TxnKV 1PC (D237).

Published figures agree on the ordering:

  • Async commit and 1PC (TiDB 5.0 release notes): Sysbench update-index average latency fell 41.7% (12.04 → 7.01 ms), and oltp-insert latency fell 37.3%. A single-region transaction skips the second phase altogether.
  • Atomic-mode raw writes (raw_compare_and_swap_atomic, raw_batch_put_atomic) run as scheduler commands that take per-key latches (tikv/src/storage/txn/commands/atomic_store.rs). TiKV's docs say that "write operations like put or delete in atomic mode are more expensive" (tikv.org/docs/7.1/develop/rawkv/cas).
  • RawKV on API v2 takes a causal timestamp from a TiKV-local TSO cache, refreshed every 100 ms, plus a concurrency-manager guard (tikv/src/storage/mod.rs). It avoids the PD round trip that a transaction's start_ts needs, but not the fence problem.
  • No current TiKV page gives RawKV put p99 for WAL-sized values. The v6.1 performance overview gives throughput only: 43.2 K updates/s on YCSB workload A on 3 × 40 vCPU, and "at most 200,000 OPS within 10 ms latency" (tikv.org/docs/6.1/deploy/performance/overview). Q114 measures it.

6.5 Key layout (D238)

  • Keyspace loams_pgwal, pre-allocated like the others (deploy/tikv/pd.toml). It keeps the WAL out of the metastore's GC safe point and placement, and gives the WAL its own quotas.
  • Keys per timeline tl = (tenant_id, timeline_id), 32 bytes:
    • H/<tl>: the head (postcard-encoded): term, term_history, flush_lsn, commit_lsn, peer_horizon_lsn, backup_lsn (bucket), remote_consistent_lsn, pg_version, system_id, wal_seg_size, timeline_start_lsn, version.
    • W/<tl>/<begin_lsn: u64 BE>: WAL bytes [begin_lsn, end_lsn) of one chunk, at most 128 KiB (one walproposer MAX_SEND_SIZE).
  • One region range per timeline. Timeline creation asks PD to split at H/<tl> and at the next timeline's prefix. The head and the tail chunks then usually share a region, so the append is 1PC. When a size split separates them, the commit falls back to async commit: prewrites go to both regions in parallel, and the ack comes after the prewrites, so latency rises by little (estimate).
  • Hotspots. Each timeline's tail is a sequential write hotspot, as a Raft log is. A timeline cannot be spread over regions without breaking ordering, and it does not need to be: one region leader handles tens of MB/s (verify). Many timelines spread across stores through PD's leader and region balancing.
  • Splits. Load-based split does not help here. It counts read QPS, bytes and CPU only, over 10 s (split.qps-threshold 3000, byte-threshold 30 MiB/s; docs "Load Base Split"), so it never splits a write-hot WAL region. The per-timeline pre-split above is what spreads the load. Size splits happen at region-split-size (256 MiB since v8.4). PD merges small regions after split-merge-interval (1 h), and the trimmed timelines leave small regions for it to merge.
  • Batching. While one append transaction is in flight for a timeline, arriving AppendRequests queue. The next transaction takes the whole queue, up to 1 MiB (8 chunks). That fits raft-max-size-per-msg (1 MiB) and is well under raft-entry-max-size (8 MiB). This is the safekeeper's own "fsync when the queue drains" group commit. With 5 ms per transaction and 1 MiB per batch, one timeline tops out at about 200 MiB/s (estimate). Pipelining more than one in-flight transaction per timeline is P4 optimisation work (Q115).
  • Values. WAL compresses well, but Postgres' WAL may already hold compressed full-page images (wal_compression). Chunks are stored as they arrive, and the bucket copy is compressed (§6.7).

6.6 TiKV tuning and the latency budget

Configuration of the loams_pgwal cluster. The WAL runs on the same TiKV release as the rest of Loams (D126). When the WAL shares the metastore's cluster, it gets its own stores through placement rules (Q116). The defaults below are TiKV v8.5's, read from pingcap/docs release-8.5 (tikv-configuration-file.md).

Settingv8.5 defaultFor the WALWhy
Raft log fsyncAlways. raftstore.sync-log was removed in v5.0 ("forcibly set to true")UnchangedA write is acknowledged only after a quorum has persisted it. Never relaxed
raft-engine.enabletrue (since v6.1). enable-log-recycle = trueUnchanged. raft-engine.dir on its own NVMe drive with power-loss protectionAn append-only multi-Raft log. PingCAP reports 20% lower tail latency and 25–40% less write IO than RocksDB (Raft Engine blog). Log recycling "reduces the long tail latency on write workloads"
raftstore.store-io-pool-size1 (async IO; 0 before v8.0)1–2. Keep store writer and raftstore threads under 80% CPUWithout it, the raftstore thread fsyncs, which "increases commit duration by about one IO+fsync time" (tikv#10540)
raftstore.store-pool-size / apply-pool-size2 / 2Raise when propose wait or apply wait growsThe tikv_raftstore_store_wf_* waterfall shows which stage waits (docs "Latency breakdown": async write = propose + commit + apply_wait + apply_log, and with async IO commit = max(wait by write worker, replicate log))
raftstore.store-max-batch-size / apply-max-batch-size, raft-write-size-limit256 / 256, 1 MiBUnchangedBatching inside TiKV already matches the WAL's 1 MiB batches
server.grpc-concurrency / grpc-raft-conn-numgrpc-raft-conn-num × 3 + 2 / max(1, min(4, cores/8)) (v8.5.4+; 5 / 1 before)Unchanged at firstClient and Raft RPC threads
raft-base-tick-interval, raft-heartbeat-ticks, raft-election-timeout-ticks1 s, 2, 10: an election after about 10 sShorter ticks for WAL stores, for example 200 ms (about 2 s to elect), with hibernate-regions (default on) to bound the CPU (Q117)A crashed leader stalls every timeline it leads until the election. Safekeepers do not stall when one of three fails. This is the WAL's weakest point (§12 row 10)
causal-ts.* (API v2)renew-interval 100 ms, alloc-ahead-buffer 3 sRaise alloc-ahead-buffer and the renew batch sizes as the docs advise, so a PD leader failover does not show as a write-latency spike ("about 15%" QPS drop otherwise)API v2 keyspaces need the causal-timestamp cache
Storage—NVMe with power-loss protectionfdatasync of 16 KB takes 0.7–10 µs on PLP drives (Intel D7-P5520, Samsung PM9A3) and 0.45–2.8 ms on consumer drives (Crucial T500, Samsung 990 Pro) (Small Datum, 2026-01). The spike ran on a consumer drive, which is why both paths measured in milliseconds
PlacementPlacement rules on (enable-placement-rules)Per timeline range (or per AZ group, Q117): one rule role: leader, count: 1, label_constraints: zone in [<compute's AZ>], and one role: follower, count: 2 in the other AZs. Keys are the memcomparable-encoded 'x' + keyspace id bounds, in hex. Loams's control plane writes the rule when it places the computeThe client→leader hop stays in the AZ. Quorum needs one cross-AZ follower

Client side (Loams's client-rust fork).

  • No RPC batching. client-rust sends every request as its own unary gRPC call (src/store/request.rs). It has no BatchCommands stream like client-go's (max-batch-size 128, max-batch-wait-time). One append per timeline in flight is unaffected. Many timelines per WAL instance may need it (Q118).
  • No follower or stale reads, and no zone-aware routing. Every request goes to the leader (upstream PR #562). The leader placement rule is what keeps the WAL in the AZ.
  • Keep a TSO prefetch. A 1PC transaction still needs a start_ts from PD. The fork's TSO oracle batches requests (R1 plan row F1). With PD's leader in another AZ, a TSO costs one cross-AZ round trip, so the WAL pool prefetches.
  • The async-commit and 1PC fixes are already in the fork. max_commit_ts, and the 1PC fallback that commits instead of failing, are in docs/plans/r1-dependency-spike.md §(i). The WAL relies on both.

Latency budget (estimate). One commit record. The TiKV leader is in the compute's AZ and the followers are in two other AZs. The disks are NVMe with power-loss protection. AWS inter-AZ RTTs measured 0.39–2.42 ms across regions, "the vast majority" under 1 ms (bitsand.cloud), and AWS itself says "single-digit millisecond".

StepLoams WALSafekeepers ×3 (AZ a, b, c)
Compute → acceptor0.1–0.2 ms (in AZ)0.1 ms to a, and in parallel one cross-AZ RTT to b
TSO for start_ts0–0.5 ms (prefetched; one cross-AZ RTT if PD's leader is elsewhere and not prefetched)—
Fence read of the head (leader)0.2–0.4 ms—
Durable writeLeader's Raft append and fsync, in parallel with one follower's: cross-AZ RTT (0.4–1 ms) + fsync (µs) + batching, 0.6–1.5 ms; then apply, 0.1–0.5 msfdatasync on a and b; the quorum is complete when b's ack arrives: one cross-AZ RTT + fsync, 0.5–1.2 ms
Ack back0.1–0.2 msincluded
Total, p50about 1.2–3.3 msabout 0.6–1.5 ms
p99 target< 5 ms (owner target)to be measured (P4b)

The honest reading.

  • The extra steps. At equal hardware, the Loams WAL's commit path has one more read, one TSO and TiKV's apply step. The p50 is likely about 0.5–1.5 ms worse than three safekeepers.

  • What the spike showed. On this laptop:

    • the single safekeeper acked in 3.57 ms at p50 and 22 ms at p99;
    • a fenced 1PC write to a single TiKV store took 6.5–7.1 ms at p50 and 25–26 ms at p99.

    That is about 2× at p50, and close at p99, both dominated by a consumer SSD's fsync.

  • What the gate compares. The gate (§7) is on p99. There, async IO, Raft Engine and PLP disks can bring the Loams WAL level with a safekeeper whose fsync is the tail.

  • Failover. A TiKV leader crash stalls its timelines for the election timeout, where losing one of three safekeepers stalls nothing. The P4b fault run measures commit latency through a leader kill as well as steady state (§7).

Where Loams wins without argument:

  • Operations. A stateless pool on a TiKV cluster that Loams already runs, instead of a stateful three-node safekeeper fleet per region with membership changes, eviction and a scheduler.
  • Durability. The bucket copy lags by the group-commit interval (250 ms). Safekeepers upload full 16 MiB segments, and partial segments only every 15 minutes. After a total loss of the hot tier, the RPO is 250 ms instead of up to 16 MiB or 15 minutes of WAL.
  • Cost. Idle timelines cost nothing in the WAL tier, because their keys are trimmed. Safekeeper nodes are provisioned for peak, with three NVMe volumes each.
  • Throughput. Group commit is per timeline, and TiKV scales out across stores.

If the gate fails on p99, the WAL stays behind its feature and stock safekeepers remain the default (D233). Nothing else in this document depends on the WAL.

6.7 How the pageserver and computes read it; retention; recovery (D240)

  • Replication for readers. The WAL service implements START_REPLICATION [PHYSICAL] X/Y [(term='N')] like send_wal.rs:
    • Without a term, it serves up to commit_lsn, for the pageserver and replicas.
    • With a term, it serves up to flush_lsn within that term, for the compute's neon_walreader and recovery.
    • The pageserver hard-codes the interpreted protocol (timeline.rs launch_wal_receiver: protobuf, zstd level 1, per-shard filtering). The WAL service therefore decodes WAL with Neon's wal_decoder crate (Apache-2.0).
    • Confirmed in P4a: vanilla ingestion was removed from the pageserver (neon #12126). A compute on the WAL service without an interpreted sender stalls on its first read of a page it wrote, because the pageserver never gets the WAL. Until the WAL service links wal_decoder, loams-safekeeper's feeder acts as walproposer towards one stock safekeeper and streams it the committed WAL, off the commit path. That safekeeper runs with --no-sync, publishes to the broker and serves the pageserver. This is interim (Q112).
  • Sources, in order:
    1. the in-memory tail of recently appended chunks (written through at append time);
    2. TiKV (scan W/<tl>/begin..);
    3. the Loams log in the bucket for ranges already trimmed from TiKV.
  • Feedback. The pageserver's 'z' updates carry remote_consistent_lsn. They are written into the head, coalesced to at most one write per second per timeline, and returned to the compute in AppendResponse.pageserver_feedback, which drives walproposer's backpressure (max_replication_*_lag). Hot-standby feedback is passed through the same way.
  • Discovery. Every instance publishes SafekeeperTimelineInfo for the timelines it serves to the unmodified storage broker, every second: safekeeper_id = the pool's logical id, safekeeper_connstr = the pool's Service, and commit_lsn, flush_lsn, remote_consistent_lsn and availability_zone from the head. It also answers SafekeeperDiscoveryRequest for any timeline in loams_pgwal.
  • Group commit to the bucket.
    • A per-timeline task appends committed chunks to the Loams log (loams-log standard class, D72's idempotent producer = (timeline, term), sequence = begin_lsn), with a 250 ms flush interval.
    • After each acknowledged log append, it advances head.backup_lsn.
    • The stream is pgwal/<tenant> with one partition per timeline slot (Q119), so retention, segmenting, encryption (D96) and GC come from the log unchanged.
  • Trimming. Chunks below min(backup_lsn, remote_consistent_lsn, commit_lsn) are deleted from TiKV in batched transactions. MVCC GC for the keyspace (loams-tikv gc.rs) reclaims them. A timeline with no compute and no lag holds no TiKV keys except its head.
  • Recovery.
    • An instance crash is a reconnect (§6.3).
    • A TiKV leader failure is a Raft election (the election timeout). Appends stall and then resume, and nothing acknowledged is lost.
    • Losing a TiKV AZ keeps quorum.
    • Losing TiKV entirely loses at most the last group-commit interval of acknowledged WAL. The timeline is recovered from the bucket's copy up to backup_lsn, and the compute must be restarted from that LSN (recover_from_bucket, an operator action). With safekeepers, the same event loses everything since the last uploaded segment.

6.8 The public Rust API (D235)

A trait-level sketch, for P4a to refine:

/// A timeline: (tenant, timeline) in Neon's ids.
pub struct TimelineId { pub tenant: [u8; 16], pub timeline: [u8; 16] }
pub type Lsn = u64;
pub type Term = u64;

/// The acceptor's durable state: what a safekeeper keeps in its control file.
pub struct AcceptorState {
    pub term: Term,
    pub term_history: Vec<(Term, Lsn)>,
    pub flush_lsn: Lsn,
    pub commit_lsn: Lsn,
    pub peer_horizon_lsn: Lsn,
    pub backup_lsn: Lsn,
    pub remote_consistent_lsn: Lsn,
    pub timeline_start_lsn: Lsn,
    pub server: ServerInfo, // pg_version, system_id, wal_seg_size
}

pub enum Fenced { Term { current: Term } }

/// The storage behind the acceptor. TiKV implements it (P4a); an in-memory
/// implementation drives the protocol tests and the TLA+-derived traces.
#[async_trait]
pub trait WalStore: Send + Sync + 'static {
    async fn create(&self, tl: &TimelineId, server: ServerInfo, start: Lsn) -> Result<AcceptorState>;
    async fn load(&self, tl: &TimelineId) -> Result<Option<AcceptorState>>;
    /// Grant a vote if `term > state.term`; durable before returning.
    async fn vote(&self, tl: &TimelineId, term: Term) -> Result<(bool, AcceptorState)>;
    /// Adopt the elected proposer's history; truncate above `start`.
    async fn elected(&self, tl: &TimelineId, term: Term, start: Lsn,
                     history: Vec<(Term, Lsn)>) -> Result<Result<AcceptorState, Fenced>>;
    /// Fenced, contiguous append of `[begin, begin + wal.len())`; durable on Ok.
    async fn append(&self, tl: &TimelineId, term: Term, begin: Lsn, wal: &[Bytes],
                    commit_lsn: Lsn, truncate_lsn: Lsn) -> Result<Result<AcceptorState, Fenced>>;
    /// Read WAL `[from, to)` from the tail cache, TiKV or the bucket.
    fn read(&self, tl: &TimelineId, from: Lsn, to: Lsn) -> BoxStream<'static, Result<Bytes>>;
    async fn record_feedback(&self, tl: &TimelineId, fb: PageserverFeedback) -> Result<()>;
    /// Delete WAL below `lsn` once `backup_lsn` and `remote_consistent_lsn` have passed it.
    async fn trim(&self, tl: &TimelineId, lsn: Lsn) -> Result<()>;
}

/// The service: the safekeeper wire protocol over a `WalStore`.
pub struct WalService<S: WalStore> { /* … */ }
impl<S: WalStore> WalService<S> {
    pub fn new(store: S, broker: BrokerClient, config: WalServiceConfig) -> Self;
    /// `START_WAL_PUSH` (proposer) and `START_REPLICATION` (readers) on one listener,
    /// plus `IDENTIFY_SYSTEM` and `TIMELINE_STATUS`, over the Postgres protocol.
    pub async fn serve(self, listener: TcpListener, shutdown: CancellationToken) -> Result<()>;
}
  • Crate layout.
    • proto: the message codecs for v2 and v3. It is pure Rust, with no Postgres headers.
    • acceptor: the state machine, ported from neon/safekeeper/src/safekeeper.rs with its tests.
    • tikv: the WalStore implementation.
    • send: readers, and the interpreted sender behind the feature interpreted.
    • broker: the storage-broker gRPC client, built from storage_broker/proto/broker.proto.
  • The interpreted sender depends on the fork's wal_decoder, postgres_ffi and utils, as a git dependency pinned to a fork revision. postgres_ffi's build.rs runs bindgen against Postgres server headers (POSTGRES_INSTALL_DIR/v1x/include/postgresql/server), so that feature needs the fork's Postgres headers at build time. CI builds them once per fork pin and caches them. The default Loams build does not enable the feature (Q112).
  • Attribution. Ported code keeps Neon's Apache-2.0 notice in NOTICE.

6.9 Mapping the safekeeper protocol to Loams operations

Message (direction)Fields that matterLoams operation
START_WAL_PUSH (proto_version '3', allow_timeline_creation …) (P→A)Query on a replication connection with tenant_id and timeline_id in optionsOpen the session; WalStore::load (or create if allowed and absent)
ProposerGreeting 'g' (P→A)tenant_id, timeline_id, mconf (generation, members), pg_version, system_id, wal_seg_sizeValidate against the head (version, segment size, system id); the membership is fixed at generation 1 with one member
AcceptorGreeting 'g' (A→P)node_id, mconf, termFrom the head: term; node_id = the pool's logical id
VoteRequest 'v' (P→A)generation, termWalStore::vote: one transaction; read the head, write term if higher
VoteResponse 'v' (A→P)term, vote_given, flush_lsn, truncate_lsn, term_historyFrom the head after the vote
ProposerElected 'e' (P→A)term, start_streaming_at, term_historyWalStore::elected: check find_highest_common_point; rewrite the chunk containing start_streaming_at to end there and delete later chunks (never below commit_lsn); write the history, in one transaction
AppendRequest 'a' (P→A)term, begin_lsn, end_lsn, commit_lsn, truncate_lsn, WAL bytesQueue. The group-commit task takes the queue into one WalStore::append (§6.4). An empty request is a heartbeat: it updates commit_lsn in the next batch or in a coalesced head write
AppendResponse 'a' (A→P)term, flush_lsn, commit_lsn, hs_feedback, pageserver_feedbackSent after the append transaction commits, and unsolicited when new pageserver feedback arrives (as network_write does)
START_REPLICATION PHYSICAL X/Y (reader)Options protocol = interpreted, shard_*, availability_zoneStream from WalStore::read up to commit_lsn (interpreted: decode and filter per shard)
START_REPLICATION … (term='N') (compute neon_walreader)termStream up to flush_lsn while the head's term equals N
'r' StandbyReply, 'h' HotStandbyFeedback, 'z' NEON_STATUS_UPDATE (reader→A)Write/flush/apply LSNs; xmin; PageserverFeedback (remote_consistent_lsn, …)record_feedback, coalesced; relayed to the proposer
TIMELINE_STATUS, IDENTIFY_SYSTEM—From the head
Safekeeper HTTP API: POST /v1/tenant/timelinemconf, pg_version, start_lsnWalStore::create. It is called by Loams's control plane, not by the storage controller
Storage broker PublishSafekeeperInfo, SafekeeperDiscoveryRequestcommit_lsn, flush_lsn, remote_consistent_lsn, safekeeper_connstr, availability_zonePublish every second from the head; answer discovery for any timeline in loams_pgwal
term_bump, membership, pull_timeline, snapshot, partial backup, eviction—Not needed with one logical acceptor (§6.3). term_bump is kept for operators as a head write

6.10 The upgrade path

  1. P2: the fork runs with stock safekeepers under Loams's control plane (x/….wal = safekeepers).
  2. P4a–P4b: the WAL service runs beside them on the benchmark topology only.
  3. P4c: a database is switched by setting wal = loams in its x/ record. The switch runs:
    • At a quiesced point: stop the compute and wait until the safekeepers' commit_lsn = flush_lsn.
    • Loams's control plane creates the timeline in loams_pgwal at that LSN (start_lsn, with the term history copied from the safekeepers' TIMELINE_STATUS).
    • The compute spec's safekeeper_connstrings becomes the WAL pool's Service, and the compute starts: sync-safekeepers runs against one acceptor, then the basebackup.
    • The safekeeper timeline is deleted after the pageserver's remote_consistent_lsn passes the switch point.
  4. Rollback is the same procedure in reverse.
  5. Branches inherit the parent database's setting.

7. The benchmark gate (P4b, D233)

The gate. On the same hardware and topology, with pgbench against stock safekeepers ×3 versus the Loams WAL:

  • p99 commit latency (Loams) ≤ p99 commit latency (safekeepers), and
  • throughput (Loams) ≥ throughput (safekeepers), for every workload below, within the run-to-run noise measured by repeating the baseline 3 times. P4c merges only when both hold.

Topology (deploy/loams-pg-bench/: a compose file for a single host, and a Kubernetes manifest for three nodes):

TierBaselineCandidate
Compute1 × fork compute image, pinned shared_buffers and max_connections, fsync = onsame image and settings
WAL3 × safekeeper (fork image), each on its own node or volumeWAL pool ×2 (one per node) + TiKV ×3 (Raft Engine, async IO, the leader rule of §6.6) + PD ×1 (or ×3)
Storagepageserver ×1, storage broker, RustFSsame
Clientpgbench in a separate container on the compute's nodesame

On a single host, tc netem adds a fixed delay between containers to model AZs: 0.5 ms one way for "cross-AZ" pairs. Both runs use the same delays.

Workloads (scripts/loams-pg-bench/):

NameScriptClientsWhy
commit-1INSERT INTO t(v) VALUES (repeat('x',100)), one per transaction1The pure commit round trip
commit-16same16Group commit
tpcb-16built-in TPC-B, -s 5016A realistic OLTP mix
tpcb-64same64Saturation, which decides the throughput criterion
bulkOne transaction inserting about 1 GB1Sustained WAL throughput (MB/s); counts toward pass or fail
bulk-burstOne transaction inserting about 250 MB1A burst, reported only; it does not count toward pass or fail

The bulk gate is a sustained write (owner decision, 2026-10-01) because a 250 MB burst measures the drive cache: on the same disk the stock safekeepers reach 118 to 182 MB/s on 250 MB and 22 to 26 MB/s on 1 GB.

Each run lasts 5 minutes after a 1-minute warm-up, with pgbench -l per-transaction logs. p50, p90, p99 and p99.9 come from the logs, not from pgbench's averages. The spike's scripts, pgb.sh and pgb-async.sh, are the starting point.

Fault run (reported, not part of the pass/fail gate unless the owner makes it one): commit-16 for 5 minutes while one TiKV leader (or one safekeeper, for the baseline) is killed at minute 2. It records the longest commit stall and the p99.9.

Recording. Each run writes bench/results/<date>-<git-sha>-<variant>.json, containing:

  • the topology and versions (fork sha, TiKV version, kernel, disk model);
  • the settings;
  • the percentiles and TPS per workload;
  • the TiKV latency-breakdown metrics (propose, append, commit and apply waits) and the safekeeper flush_wal histograms.

A summary table goes into the P4c PR description. The run is a manual workflow (workflow_dispatch) on a dedicated runner, because shared CI runners are too noisy for p99s. It is also runnable locally. The spike's single-host numbers (§6.1, §6.4) are the first data point, not the gate.

7.1 First results: laptop, 2026-09-29 (P4a/P4b; not gate data)

Implemented: crates/loams-safekeeper:

  • the v3 codecs, the acceptor and WalStore;
  • the TiKV store, as fenced 1PC transactions;
  • the service, HTTP API, feeder and the loams-wal binary.

The harness is in deploy/loams-pg-bench and scripts/loams-pg-bench, and the raw results in bench/results/.

Verified: an unmodified Neon compute (compute-node-v16 at 77e22e4):

  • starts through sync-safekeepers against loams-wal on TiKV v8.5.8;
  • commits through it;
  • is ingested by the pageserver through the feeder.

Setup:

  • Host: one laptop, 14 cores, 15 GB of RAM, a consumer Samsung NVMe without power-loss protection, btrfs.
  • Placement: every tier on the same disk, with no injected cross-AZ delay.
  • Other load: other work shared the host.
  • Runs: 60 s per workload after a 10 s warm-up. Baseline and candidate were interleaved three times each.
  • Compute settings: shared_buffers = 2GB, compute fsync = off.

Means over the three repeats:

WorkloadrfSafekeepers p50 / p99 (ms)Safekeepers TPSLoams p50 / p99 (ms)Loams TPS
commit-112.9 / 20.01996.5 / 43.1109
commit-1616.0 / 28.91 96917.0 / 66.3791
tpcb-16 (-s 10)115.7 / 12568728.8 / 237365
bulk (WAL MB/s)1–97.5–20.4
commit-135.6 / 37.111713.8 / 96.258
commit-16311.5 / 63.698332.4 / 127474
tpcb-16 (-s 10)329.0 / 30036956.1 / 443215
bulk (WAL MB/s)3–76.9–11.9

rf 1 compares one safekeeper with one TiKV store; rf 3 compares three safekeepers with three stores.

These are not gate results. They are laptop data points outside the §7 protocol: one host, a consumer disk, 60 s runs instead of 5 minutes, -s 10, and no tpcb-64. The gate itself needs the §7 topology on a dedicated runner (the manual loams-pg-bench workflow).

On this laptop, the comparison script's criteria fail for every workload at both replication factors. The Loams WAL's p99 is about 1.5–2.6× the safekeepers', and its throughput about 0.2–0.6× (bench/results/gate-rf*.md). This gap is far wider than the noise, so a §7 run is not worth doing until the levers below are in. Meanwhile, D233 applies: the WAL stays behind its feature, and P4c does not start.

Where the time goes. TiKV's own metrics, from a commit-1 run:

  • Raft log persist dominates: p50 ≤ 2.6 ms, p99 up to 82 ms. That is the same consumer-SSD fsync the safekeeper pays, but with TiKV's scheduler, apply and 1PC steps added, plus three client RPCs per append: TSO, the fenced head read and the prewrite.
  • Single-flight appends per timeline: they cap group commit. bulk shows it most, at 12–20 MB/s against 77–98 MB/s.
  • The interim feeder adds a second WAL stream on the same disk.

Before the gate can pass, the levers are:

  1. Pipelined appends (Q115). Several appends in flight per timeline, fenced without serialising on the head key. This is the main lever for commit-16, tpcb and bulk.
  2. No PD round trip per append. Reuse the previous commit timestamp as the next start_ts. This is safe here: a stale start_ts only causes extra write conflicts. It needs a begin_at(ts) in the client-rust fork.
  3. The in-process interpreted sender (Q112). It removes the feeder's second stream.
  4. Re-run on the §7 topology. Server NVMe with PLP (fsync in µs, not ms), three nodes and real or netem AZ delays (Q114). The owner's < 5 ms p99 target is a server-hardware number. Both variants on this laptop are far above it.

7.2 Arm A: a native leaderless WAL on local NVMe (D263–D269 and D271–D272, 2026-09-30)

On 2026-09-30 the owner decided that the Loams WAL targets native leaderless-WAL performance instead of tuning the TiKV hot tier. Two arms are measured against the same gate:

  • Arm A, this section, takes TiKV off the commit path.
  • Arm B keeps TiKV and removes its per-append overheads.

The owner also asked for Arm A's store to be a tiered, low-level I/O layer, with runtime detection and fallback (D266), and chose compio as its runtime (D263). The decisions below are Loams's proposals within that direction. They supersede D239 (one logical acceptor per timeline) for Arm A only, and leave D237–D238 in force for Arm B.

The causes in §7.1, and how Arm A removes each one:

§7.1 causeArm A
TiKV's Raft hop and apply on every appendNo TiKV on the commit path. Each acceptor writes to its own local NVMe
Three client round trips per append (TSO, fenced head read, prewrite)None. The term fence is an in-memory check on the acceptor, made durable by the vote
One append in flight per timelineMany in flight. Writes are issued while earlier ones are still syncing, and acks follow the durable position (Q115; adoption awaits P4b results)
The feeder shares the diskThe feeder stays until Q112 is solved, but its safekeeper moves to another filesystem; a separate device is needed to remove device contention (D271)

The design (target; D272 records the compio path as built, which leaves out the custom OpCodes and registered buffers below)

  • D263. compio is the runtime of the Arm A data path (owner decision, 2026-09-30; final). The Arm A build enables compio's io-uring and polling features.

    • Shards. Each shard is one thread running a compio runtime, and owns a set of timelines (by hash of the timeline id, over a shard count fixed at first start). A shard also owns its own journal (D265), so group commit is per shard and needs no locks across cores.
    • What runs on the owning shard: the timeline's START_WAL_PUSH connection, its journal writes, and its durable writes. An accept thread reads the startup packet, which names the timeline, then hands the socket to the owning shard.
    • The driver is io_uring. (Target design. The built path in D272 uses O_DSYNC descriptors and no custom OpCodes or registered buffers.) compio's own opcodes cover sockets. Custom OpCodes cover what compio does not expose: WRITE/WRITE_FIXED with RWF_DSYNC, and registered buffers. The raw io-uring crate and io_uring_register on the ring's fd are used only where compio cannot help: buffer registration and file-table updates.
    • Linked writes. The safe compio 0.19 API used here cannot pass the IO_LINK flag on an operation. So the WRITE → FSYNC pair for FUA-less devices is two back-to-back operations on the shard, and the FSYNC is submitted only after the WRITE completes successfully. An AppendResponse never precedes durability. (The low-level Extra / Proactor::push_with_extra path could link the pair; it is a later optimisation.)
    • When the polling feature is enabled and the driver selects it, compio uses its epoll driver when io_uring is unavailable; durable writes use the pwritev2 pool.
    • tokio stays for the control plane: the TiKV metadata client, S3 offload, the admin HTTP API (axum) and the feeder. It is bridged to the shards by bounded channels, off the commit path. The existing tokio WAL service stays as the fallback front end (--runtime tokio), and is benchmarked as tier (a).
    CrateLicenseLatest release (crates.io, 2026-09-30)ModelRole
    compioMIT0.19.2 stable, 2026-08-18 (0.20.0-beta.1, 2026-09-29)Thread-per-core, completion-based (io_uring, epoll fallback)The data path (D263). Apache Iggy moved to it (verify)
    io-uring (tokio-rs)MIT OR Apache-2.00.7.15, 2026-09-07; already in Cargo.lock through qdrant-edge and compioRaw ringsOnly where compio does not reach
    monoio (ByteDance)MIT OR Apache-2.00.2.4, 2024-08-20Thread-per-coreNot used. No release for over two years; dropped from the benchmark (owner)
    glommio (Datadog)Apache-2.0 OR MIT0.9.0, 2024-03-25Thread-per-coreNot used. No release for over two years
    tokio-uringMIT0.5.0, 2024-05-27io_uring under tokioNot used. No release for over two years
  • D264. walproposer is the only sequencer.

    • The compute's neon.safekeepers names three loams-wal acceptors, and walproposer runs its own Paxos over them, exactly as over safekeepers: terms, votes, ProposerElected, and commitLsn as the quorum of flush_lsn.
    • Commit is 1 round trip and 1 durable write: a majority of acceptors has the WAL on local disk.
    • An acceptor's term check is in memory. It is raised only after the vote is durable (D268), which is the Paxos acceptor's promise.
    • The vote handler first waits until the timeline's written WAL is durable. So VoteResponse.flush_lsn is what the disk holds, as in Neon's handle_vote_request.
    • After the vote, appends of a lower term are refused, and any later response carries the higher term, so a deposed proposer steps down.
    • Losing one of the three acceptors stalls nothing, which removes §12 row 10's TiKV leader-election stall for Arm A.
  • D265. One shared journal per shard, for every timeline the shard owns.

    • Precedent: TiKV's raft-engine, which puts all regions in one log.
    • Segments. The journal is a directory of fixed-size segment files, 64 MiB by default. A background thread prepares the next two:
      • it fallocates each one, then pre-zeroes it with large direct writes and one fdatasync, so the first append never converts an unwritten extent (a metadata journal commit);
      • after that, segments are recycled instead of created.
      • On btrfs the directory gets NOCOW (FS_NOCOW_FL, as systemd-journald does), so overwrites stay in place.
    • Records are framed: kind, timeline, term, LSN, length, and a CRC32C seeded with the segment's sequence number, so a recycled segment's old records never validate. The kinds are:
      • Append: WAL bytes, plus the proposer's commit_lsn.
      • Truncate: a logical truncation of one timeline, on election.
      • Progress: commit_lsn, backup_lsn and remote_consistent_lsn learned off the append path.
    • Flush units. Records are written in flush units aligned to 4 KiB, zero-padded to the block size. This is what O_DIRECT needs. It also lets several units be in flight at once without two writes ever touching the same block. The cost is space: at most one block per unit, on WAL that is offloaded and recycled.
    • Group commit. One writer coalesces every pending record, from every timeline and every connection on the shard, into the next unit (up to 1 MiB), and keeps up to --io-depth units in flight. The shard's durable position is the end of the longest prefix of completed units. By contrast, a safekeeper fsyncs each timeline's segment file separately, and writes nothing while its fsync runs.
    • Recovery replays segments in order and stops at the first record that fails its CRC (the torn tail). Units complete out of order, but an ack only covers a prefix of completed units, so nothing acknowledged is behind a torn unit. Each timeline keeps only the WAL contiguous from its trimmed point, and its flush_lsn is where that WAL ends.
    • Reads. An in-memory index maps each timeline's LSNs to records. Reads are preads.
  • D266. A tiered I/O layer, chosen at startup. loams-wal probes the journal directory and its block device, logs the tier it picked, and takes --io auto|uring|pwritev2|buffered as an override.

    TierDurable writeUsed when
    uring (target design; D272 is the built path)io_uring on the owning compio shard (D263), with registered files (segment slots updated on rollover) and registered, aligned buffers (WRITE_FIXED) on O_DIRECT segments. The write carries RWF_DSYNC: on pre-zeroed blocks the kernel sends one FUA write when the device has FUA, and write + flush when it does not. The alternative is a linked WRITE → FSYNC(DATASYNC) pair (--uring-sync linked, chosen automatically when queue/fua is 0). Optionally SQPOLL (--uring-sqpoll). IOPOLL is used only when the device has NVMe poll queues (queue/io_poll = 1)io_uring is allowed (io_uring_setup succeeds and the opcodes probe) and the filesystem takes O_DIRECT
    pwritev2O_DIRECT + pwritev2(RWF_DSYNC) from a small thread pool (one thread per in-flight unit)io_uring is blocked, as by Docker's default seccomp profile and GKE Autopilot's containerd RuntimeDefault seccomp profile
    bufferedpwrite into the page cache; a sync thread runs fdatasync back to back while unsynced data existsNo O_DIRECT (tmpfs, some overlay filesystems), or chosen by --io buffered
    Future: NVMe passthroughuring_cmd on /dev/ng*: NVMe write commands with FUA, no filesystemA raw namespace per node. Documented only; no SPDK
    • This drive. The laptop's drive reports fua = 1, write_cache = write back, io_poll = 0 (no poll queues), and 512-byte logical blocks.
    • Network. Every connection sets TCP_NODELAY. SO_BUSY_POLL can be set with --busy-poll-us, and is logged and ignored when the kernel refuses it.
    • io_uring networking. The shard's sockets go through compio's io_uring driver. Multishot recv with provided buffer rings (compio's buffer pool) is the next step (Q264).
    • kTLS is the cheap route for issue #148: a rustls handshake, then TLS_TX/TLS_RX offload, so the data path keeps plain writes. That is a note only, not built.
  • D267. Pipelined appends.

    • The service writes each group of queued AppendRequests as soon as it arrives, and keeps reading.
    • An AppendResponse goes out whenever the durable position passes some of the timeline's writes. It carries the highest durable flush_lsn.
    • Heartbeats are answered at once with the durable position.
    • Stores that are durable per call (memory, TiKV) keep the old behaviour through the trait's default method.
  • D268. TiKV holds only metadata.

    • What: per (node, timeline), the acceptor record: term, term history, membership, server info, the start LSNs and trimmed_lsn.
    • When: it is written on timeline creation, vote, ProposerElected and trim.
    • How: these are 2PC transactions, which is fine at one or two writes per election, off the commit path.
    • Order on election: a Truncate record is made durable first, then the metadata. A crash between the two leaves a shorter WAL under the old history, which the next election repairs.
    • Development: a local control-file backend (write, fsync, rename, fsync the directory) serves single-node setups and tests.
  • D269. Offload and trim.

    • Offloader: one acceptor per timeline, chosen by a lease in TiKV (B/<tl>: owner, expiry, backup_lsn). It uploads committed WAL to the bucket every 250 ms (or at 16 MiB).
    • Objects: pgwal/<tenant>/<timeline>/<begin>-<end>.lwal. The format follows loams-log's segment conventions: a magic and version header, a CRC32C trailer, and a zstd body. The Kafka RecordBatch framing is not used, because WAL is one byte stream per timeline.
    • backup_lsn is advanced in the lease record, and the other acceptors read it from there.
    • Trim. A segment is recycled once every timeline in it has passed min(backup_lsn, remote_consistent_lsn, commit_lsn).
    • Pinning. An idle timeline whose pageserver lags pins old segments. Rewriting its live records forward, as raft-engine purges, is left for later (Q262).
  • D271. The pageserver feed stays the feeder for now (Q112).

    • Why not in-process decoding yet: it needs Neon's wal_decoder, postgres_ffi (bindgen against the fork's Postgres server headers), utils and pageserver_api, which bring Neon's workspace dependency pins into Loams's. That is not feasible within this arm.
    • Instead:
      • exactly one designated acceptor runs the feeder;
      • it feeds up to min(commit_lsn, its durable flush_lsn);
      • its stock --no-sync safekeeper keeps its data on a different filesystem from the acceptors' journals.
    • On the laptop that filesystem is on the same physical drive, so the feeder's writeback still competes for the device.

1PC, 2PC or raw: where each write goes

The commit path is the only place where latency is gated. TiKV's modes are measured in §6.4 (single store, spike). The local-disk figures are from 2026-09-30 on the same laptop: 400 writes to a 64 MiB pre-zeroed file, with another session's build running.

Write modeWork per commitClient round tripsFencep50 / p99 on the laptopArm A uses it for
TiKV TxnKV 1PC (P4a)TSO, fenced head read, prewrite, Raft append + apply3Write-write conflict on the head6.5–7.1 / 25–26 msNothing
TiKV TxnKV 2PCTSO, read, prewrite, commit, two Raft appends4Same19–26 / 57–74 msMetadata: create, vote, elected, trim (one or two writes per election)
TiKV RawKV put, or CAS + putOne (or two) Raft appends1–2None, or a CAS on the head6.4–14.9 / 33–51 ms; fenced 13.8–29.8 / 85–595 msNothing. Unfenced writes can lose acknowledged commits with one acceptor (§6.4)
Local, buffered pwrite + shared fdatasyncOne pwrite and a share of one sync, on 2 of 3 acceptors in parallel1 (compute ↔ acceptor)In memory, behind a durable votebtrfs 2.4 / 3.0 ms per sync, NOCOW 2.1 / 2.8; ext4 0.9 / 1.3The commit path, tier buffered
Local, direct O_DIRECT + RWF_DSYNC (FUA), 4 KiBOne write, FUA or write + flush1Samebtrfs 2.2–2.7 / 3.6–4.4 ms at queue depth 1; ext4 0.9 / 1.5The commit path, tiers uring and pwritev2
Same, 16 writes in flight———btrfs 4.0–17.6 / 10.7–27.7 ms; ext4 2.4 / 49.6 msNot as a default. --io-depth stays small, and one unit carries many commits

What the table says.

  • On this drive a FUA write costs about as much as a cache flush. So on equal filesystems, the direct tiers and fdatasync land close together.
  • btrfs adds its log-tree commit to every durable write: 2.1–2.7 ms here against about 0.9 ms on ext4, on the same drive.
  • A power-loss-protected drive matters more than any tier. Its durable write is microseconds (§6.6), and that cannot be shown on the laptop.
  • What the tiers remove is everything else: allocation, fsyncs and renames on the commit path, a sync that waits for earlier writes, and one sync per timeline. The gate measures each tier separately (§7.3), and also on an ext4 partition, to separate the btrfs cost.

Open questions (Q261–Q264)

#QuestionNeeded by
Q261Membership changes for Arm A: reuse Neon's generations and pull_timeline, or have the control plane write a new member set in TiKV and re-sync the new member from the bucket plus a peer?P4c
Q262Rewrite (purge) the live records of idle timelines forward, so they stop pinning old journal segments?Before multi-tenant production
Q263Tune the defaults on PLP NVMe, where the durable write is microseconds and per-syscall overhead dominates: --io-depth, SQPOLL, and IOPOLL with poll queuesThe three-node gate run
Q264Multishot recv with provided buffer rings on the shards' sockets, and SO_BUSY_POLL by default?After Q263, if network or scheduling shows in the p99

The results of the Arm A gate runs are in §7.3.

8. PgDog routing (D236)

Proposed 2026-10-01 (§31, D304–D307): sharded Loams Postgres databases use PgDog's own sharding, with the shard map in Loams's metastore rendered into pgdog.toml; cutover across several PgDog instances is orchestrated by Loams with a fence in Postgres (ALTER ROLE … NOLOGIN), because PgDog's open-source RESHARD cuts over one instance only; PgDog's 2PC is off by default and allowed only with a durable coordinator log (StatefulSet, NODE_ID, DEPLOYMENT_ID, the WAL directory on a volume); computes of 2PC databases get max_prepared_transactions in their spec. D236 is unchanged: PgDog stays unmodified.

Verified 2026-09-29.

  • License: pgdogdev/pgdog is AGPL-3.0 (LICENSE is the GNU AGPL v3 text; gh api repos/pgdogdev/pgdog). There is no CLA. A closed Enterprise Edition exists (control plane, query monitoring, QoS), built from a private repo (docs.pgdog.dev/enterprise_edition).
  • Activity: v0.1.60 (2026-09-24). Releases are weekly, and the last commit was 2026-09-28.
  • Features (README; docs.pgdog.dev):
    • transaction and session pooling;
    • load balancing across replicas (round robin, random, least connections) with health checks;
    • failover detection from replication state (it is "not a replacement for Patroni");
    • sharding (hash compatible with Postgres partitioning, list, range, schema);
    • cross-shard queries with partial aggregate support, two-phase commit, and online resharding over logical replication;
    • prepared statements in transaction mode;
    • SCRAM, MD5 and TLS;
    • an admin database, OpenMetrics and a Helm chart (MIT).
  • Routing by database name: each [[databases]] name is what clients connect to, with its own host, port, database_name and pool (docs.pgdog.dev/configuration/pgdog.toml/databases). That is exactly §23's D153 routing, without Loams writing a splice.

Deployment.

  • Where it runs: PgDog is a Deployment in the Loams Postgres namespace, with an image from ghcr.io/pgdogdev/pgdog pinned by digest.
  • Config: Loams's control plane renders pgdog.toml and users.toml from the x/ and X/ records into a ConfigMap and a Secret, and sends RELOAD to PgDog's admin database after each change.
    • Database names: <db> for main and <db>__<branch> for branches (this answers Q49's syntax; the options form stays unsupported).
    • Hosts: the branch's compute, plus read replicas (Neon replica computes) with role = replica.
  • Auth:
    • PgDog terminates client auth (SCRAM) against users.toml. Loams provisions the same role and password into the compute spec, and PgDog uses them for server connections.
    • PgDog's passthrough auth forces auth_method = plain (docs.pgdog.dev/features/authentication), so it is not used.
    • Credentials come from the auth plan's credential store (Q30).
    • TLS on both hops, with peer verification. Clients reach PgDog only over TLS, with a certificate for PgDog's service name that clients verify (sslmode=verify-full); PgDog's non-TLS listener is not exposed. PgDog connects to computes over TLS with full verification of the compute's certificate against the cluster CA, and the compute spec accepts only TLS connections (hostssl). Where a service mesh provides mutual TLS between the pods, that satisfies the backend hop instead. Plaintext is refused on both hops.
  • Pooling mode: transaction mode by default. Session mode is used for migrations, LISTEN/NOTIFY and advisory locks. Each database entry sets its own mode.
  • What PgDog cannot do: it cannot wake a stopped compute, so there is no scale-to-zero. Scale-to-zero needs Neon's proxy and the wake_compute API (§5.2, Q113).

AGPL risk, and the rules Loams follows.

  • Never linked or vendored. No PgDog crate appears in any Cargo.lock. deny.toml already rejects AGPL, and the rule is extended to the Loams Commons and loam-platform repos. No PgDog source is copied into any Loams repo.
  • Never modified. Loams uses upstream images and configuration only. A needed fix is reported upstream (with the owner's approval) or worked around with configuration. If neither is possible, the fallback is used.
  • Why:
    • AGPL-3.0 §13 requires anyone who modifies PgDog and lets users interact with it over a network to offer those users the modified source. Running Loams cloud's routing through PgDog is exactly "interaction over a network".
    • An unmodified PgDog carries no source obligation beyond the upstream source and license already being public. Images that Loams distributes (BYOC, the showcase's compose) must still ship the license and point to the corresponding upstream source.
    • Anyone who links or vendors PgDog into Loams's code would put that code under AGPL (D11).
  • Recorded in:
    • LICENSES.md in the Loams Commons and deploy repos;
    • the license-check CI job (D-SC-6), which checks deploy images;
    • a deploy/pgdog/README.md that states these rules.

Fallbacks.

PgDogNeon proxyPgBouncer
LicenseAGPL-3.0 (service only)Apache-2.0 (in the fork; may be modified)ISC
PoolingTransaction and sessionNone (it is a router)Transaction and session
Route by database nameYesBy endpoint (SNI or options=endpoint=) through wake_computeYes ([databases])
Replica load balancingYesNoNo
ShardingYesNoNo
Scale-to-zero (wake)NoYes, through Loams's wake_compute (§5.2)No
Loams workConfig renderingThe control-plane auth API (§5.2)Config rendering
Latestv0.1.60 (2026-09-24)the fork (last upstream proxy fix 2026-05-25)1.26.0 (2026-09-23, fixes 3 CVEs)

For production scale-to-zero, the expected shape is Neon's proxy in front, which wakes the compute and routes, and PgDog behind it per database for pooling. Q113 decides.

9. Where Loams can beat Neon beyond the WAL (P5)

Each item below is proposed. P5 plans each one separately, after P4.

IdeaWhatExpected gainCost and risk
Shared object-store cache for pageserversServe on-demand layer downloads (remote_storage reads) through an AZ-local cache tier built on loams-cache (RAM and NVMe, checksummed ranges over immutable objects) and foyer, shared by every pageserver in the AZCold-tenant attach and on-demand layer download at NVMe latency after the first reader; fewer bucket GETs. Neon's layers are immutable, which is the case the cache is built forA remote_storage backend in the fork (Neon's GenericRemoteStorage enum), with a small patch
Layer files on Loams's formatsStore image and delta layers as Loams objectsLittle: Neon's layer format is tuned for page versions and GetPage@LSN, and Loams's columnar formats are notNot recommended. Kept here because the owner asked
The storage controller on TiKVReplace its Postgres (Diesel; tables for tenant shards, nodes, safekeepers, timelines and metadata health) with a persistence trait over TiKVOne fewer database to run; the control plane and the controller share TiKVA large patch across storage_controller/src/persistence.rs that every fork rebase carries. Only after the fork is stable (Q48)
Branch per agent workspaceD155 on Loams Postgres (§15 §9)Already designed; P2 enables it—
Scale-to-zerowake_compute in Loams's control plane, and Neon's proxy (Q113)Idle databases cost only bucket bytes, which the WAL design (§6.7) makes true of the WAL tier as wellThe proxy's control-plane API (§5.2)

10. Maintaining the fork (D231, D241)

The starting point, verified 2026-09-29.

FactEvidence
dina-kar/neon created as a GitHub fork. Upstream main is at fa504217c (2026-08-31, a typo fix). About 6 commits since October 2025, against 100+ a month through July 2025gh repo fork; gh api repos/neondatabase/neon/commits
Neon's staff: "our engineering team is currently 100% focused on unifying Neon and Lakebase on Databricks infrastructure … This is happening outside of the open source repos" (Discord, 2025-09-03, quoted in Discussion #12835). No open-source roadmap sinceneondatabase/neon Discussion #12835
Neon main pins its Postgres at 16.9 / 17.5 (May 2025) and has no vendor/postgres-v18neon/vendor/revisions.json, .gitmodules
neondatabase/postgres has moved on: REL_16_STABLE_neon, REL_17_STABLE_neon and REL_18_STABLE_neon, last committed 2026-04-08/09, based on 16.12 / 17.8 / 18.2gh api repos/neondatabase/postgres/branches/…
Neon's patch size over upstream: 16 → 153 commits, 149 files, +6 325/−693; 17 → 155 commits, 149 files, +6 446/−818; 18 → 106 commits, 157 files, +7 606/−632 (compared against the merged upstream stamp)gh api repos/neondatabase/postgres/compare/<stamp>...REL_1x_STABLE_neon
neon/docs/core_changes.md lists 23 core changes. Only two serve walproposer: the backpressure hook in ProcessInterrupts, and shutting walproposer down after the checkpointerneon/docs/core_changes.md
Upstream today: 18.6 / 17.11 / 16.15 (2026-08-13, fixing 28 CVEs). 18.5 was never shipped. Next minors on 2026-11-12, 2027-02-11 and 2027-05-13. PG 14 reaches EOL on 2026-11-12. PG 18 GA was 2025-09-25postgresql.org/support/versioning, /developer/roadmap, the 2026-08-13 release note
So Neon's shipped compute is 6 minors behind on 17 (17.5 → 17.11). Even Neon's newest branches are 3 minors behind, and miss the May and August 2026 CVE fixes (at least 39)the above

The .gitmodules URLs are relative (../postgres.git), so dina-kar/neon's submodules resolve to dina-kar/postgres, which P2 must create as a fork of neondatabase/postgres (Q110).

Plan.

  1. Catch-up (P2, estimate: 2–3 engineer-weeks).

    • Fork neondatabase/postgres.
    • Move the submodules to the REL_16/17_STABLE_neon heads. These are newer than main's pin and may need extension changes that were never published (Q110).
    • Merge upstream 16.15 and 17.11.
    • Build the compute images in the fork's CI.
    • Run Neon's pg_regress and the Python test_runner subset for compute and safekeepers, plus deploy/neon's smoke.
    • Drop v14 and v15, which are out of Loams's scope.
  2. PG 18 enablement (P2 or later, estimate: 3–6 engineer-weeks). Neon main has no v18 wiring. It needs:

    • vendor/postgres-v18 from REL_18_STABLE_neon (18.2) merged to 18.6;
    • v18 in libs/postgres_ffi (bindgen, pg_constants, WAL record layouts);
    • libs/wal_decoder and the pageserver's per-version walingest paths;
    • pgxn/neon built against 18;
    • the compute Dockerfile and extensions.

    CNPG covers PG 18 for the showcase meanwhile.

  3. Steady state (D241).

    • Nightly job: try merging upstream REL_17_STABLE and REL_18_STABLE into the _neon branches. Report conflicts. Build the compute image. Run pg_regress.
    • Release day (the second Thursday of Feb/May/Aug/Nov): merge the tagged minor, run the full suite, and publish ghcr.io/dina-kar/compute-node-v17:<minor> and -v18 within a week.
    • Out-of-cycle releases follow the same path.
    • Estimate: 2–4 engineer-days per quarter for two majors, because minor releases rarely touch the areas Neon patches (smgr, WAL redo, SLRU). Plus about 1 engineer-week a year for a new major's catch-up, and the Rust toolchain and dependency upkeep of the Neon workspace (cargo deny on the fork).
  4. Owning the rest. Storage-side fixes in the pageserver, broker and controller are taken as needed, with a test per fix, like the client-rust fork (§20). The fork's README states that it is independent of Neon Inc.

11. Roadmap: phases (D233, D235)

Each phase is small stacked PRs. The P4 phases are behind the feature loams-wal and change no default path.

PhaseScopeDepends onDone when
P1CNPG for the showcase: deploy/cnpg/ (kind or k3d, operator 1.30.1, Barman Cloud plugin 0.15.0, ObjectStore on RustFS), the commons-pg Cluster, a PITR test; §22's compose keeps plain postgres:17.11—PITR restores to a timestamp between two writes; OpenFGA's and GlitchTip's migrations run
P2aFork upkeep: dina-kar/postgres, submodule move, 16.15/17.11 merge, fork CI and images (§10 step 1)Q110The fork's compute images pass pg_regress and deploy/neon's smoke
P2bThe Loams control plane (§5) driving stock safekeepers: the spec endpoint, notify-attach and notify-safekeepers, the C/ records, the storage controller with its database on CNPGP2a; §23 N1–N2A compute started by Loams serves a branch; a pageserver migration pushes a new spec
P3PgDog routing: config rendering, RELOAD, the branch-name scheme, deploy/pgdog/ with the license rulesP2bpsql, psycopg 3, node-postgres and JDBC reach main and a branch through PgDog; transaction pooling holds
P4aloams-safekeeper: codecs, acceptor, WalStore, the TiKV backend, readers and the interpreted sender, broker publishing; protocol tests from Neon's safekeeper tests and TLA+ traces; fault tests (instance kill mid-append, TiKV leader kill, term bump race)P2b; Q112A compute runs against the WAL service in deploy/neon; the pageserver ingests; no acknowledged commit is lost under the fault set
P4bThe benchmark harness (§7): deploy/loams-pg-bench/, scripts/loams-pg-bench/, result files, the manual workflowP4aBaseline repeated 3 times; noise band recorded
P4cThe switch-over (§6.10) behind loams-wal, per databaseP4b gate passesMerged only with a results table that meets §7
P5Pageserver efficiency (§9): the shared cache first, then scale-to-zero (Q113), then the controller on TiKVP4c (or P2b for the cache)Per item

12. Risks

#RiskMitigation
1The Loams WAL misses the p99 gate (§6.6: one more read and one TSO than a safekeeper)D233: it stays behind the feature, and safekeepers stay. Tuning (async IO, leader placement, TSO prefetch, PLP NVMe) comes first. Pipelined appends (Q115) come next
2Protocol subtlety: terms, truncation and the basebackup-LSN PANIC path in walproposerPort safekeeper.rs's state machine with its tests. Replay Neon's TLA+ model (safekeeper/spec/) as trace tests. Run Neon's Python safekeeper tests against the service in the fork's CI
3Building wal_decoder needs Postgres headers in the Loams buildFeature-gated (interpreted); headers cached per fork pin in CI (Q112)
4TiKV becomes the WAL's single point of failure for every timelineThree AZs; the WAL gets its own stores through placement rules (Q116); the bucket copy bounds RPO at 250 ms if TiKV is lost
5TiKV MVCC and compaction churn from short-lived WAL keysA dedicated keyspace with its own GC safe point; batched deletes; monitor RocksDB write amplification; Titan for large values if needed (verify)
6Fork upkeep outgrows the estimate (PG 18 enablement; hidden dependencies between the _neon branches and unpublished extension code)§10 steps 1–2 first, estimated separately. CNPG stays the showcase's Postgres regardless (D230)
7AGPL exposure if someone patches or links PgDog§8 rules, deny.toml, the license-check job; Neon proxy and PgBouncer as fallbacks
8RustFS with the Barman Cloud plugin fails on SigV4 checksum headersThe two AWS_*_CHECKSUM_* variables; P1 tests it first
9Leader placement per timeline costs PD rules at scale (one rule per timeline)Group timelines by compute AZ into a few key ranges per AZ, not one rule each (Q117)
10A TiKV leader crash stalls commits on its timelines until the election (about 10 s with the default ticks). Losing one of three safekeepers stalls nothingShorter Raft ticks on the WAL stores (Q117); PD leader transfer before planned restarts; P4b's fault run reports commit latency through a leader kill, and the owner decides whether that failover time is acceptable alongside the p99 gate

13. Open questions

#QuestionNeeded by
Q110Fork neondatabase/postgres as dina-kar/postgres (required by the fork's relative submodule URLs), and base the catch-up on the REL_1x_STABLE_neon heads (16.12/17.8, which may need unpublished extension changes) or on main's pins (16.9/17.5)?P2a
Q111PGroonga for Zulip on CNPG: a custom image on 17.11-standard-trixie, or a separate Cluster with an image that has it?P1
Q112Where the interpreted sender builds: in Loams behind a feature, with the fork's Postgres headers in CI, or as a small binary crate inside the fork's workspace that links loams-safekeeper's WalStore? P4a: interim answer is the feeder (§6.7); the published neon image ships Postgres 14–17 server headers under /usr/local/v1x/include, so CI can extract them for postgres_ffiP4a plan
Q113Scale-to-zero: run Neon's proxy (with Loams's wake_compute) in front of PgDog, or accept always-on computes for Loams Postgres in the first release? (Supersedes Q47)P3 plan
Q114Re-measure RawKV and TxnKV 1PC write latency on a three-node, three-AZ TiKV cluster with PLP NVMe, and collect published TiKV p99 figures for WAL-sized valuesP4a
Q115Pipelining: more than one in-flight append per timeline (for example a separate term key with Lock mutations and a flush marker per batch), if the single-flight group commit limits throughputP4b results
Q116The WAL on the metastore's TiKV cluster (own stores through placement rules) or on a dedicated TiKV cluster?P4a plan
Q117Leader placement granularity (per timeline or per AZ group) and whether to shorten TiKV's election timeout for WAL storesP4b
Q118Add BatchCommands RPC batching to Loams's client-rust fork (client-go's max-batch-wait-time), if P4b shows RPC overheadP4b
Q119The Loams log layout for the bucket copy: one stream per tenant with a partition per timeline, or a stream per timeline (streams are cheap; partitions are fixed at creation)P4a

14. Contradictions with earlier decisions, and how they are resolved

EarlierConflictResolution
D149 (§23): Neon for the showcase appsD230: CNPGSuperseded (owner, 2026-09-29). Q45 is answered "no". D-SC-12's Postgres front end over TiKV stays long-term, as §23 §5 described for this case
D151 (§23): fork only when neededD231: fork nowSuperseded
D150 (§23): Loams as a client of Neon's APIsD232: Loams serves the control-plane APIAmended: it extends D150
D153 (§23): Loams's pg listener splices to computesD236: PgDog routesAmended for Postgres OLTP. The pg listener keeps analytics (PG1). Q49's syntax becomes <db>__<branch>
D11: no AGPL dependenciesPgDog is AGPL-3.0Not a dependency: an unmodified separate process, as for WeSQL (D148) and Alternator (D60)
D126: TiKV unmodifiedThe WAL tunes TiKVConfiguration only; no fork of TiKV or PD
D2 / D130: OLTP out of scopeLoams Postgres is OLTPAs in §23 §14: a separate service. D130's TiKV-backed product-line exception covers the WAL's hot tier
§22 §6.4: "one Postgres 17"CNPGConsistent: CNPG runs Postgres 17 (§22 is amended to name CNPG)
Q47, Q48 (§23)§5, §8Q47 → Q113. Q48 answered: the storage controller's database is a small CNPG Cluster; TiKV later (§9)

15. Sources

Read on 2026-09-29.

  • Neon (fork dina-kar/neon at fa504217c):
    • pgxn/neon/walproposer.c, walproposer.h, walproposer_pg.c, neon_walreader.c;
    • safekeeper/src/{safekeeper.rs, receive_wal.rs, wal_storage.rs, control_file.rs, send_wal.rs, send_interpreted_wal.rs, wal_backup.rs, wal_backup_partial.rs, remove_wal.rs, timeline_manager.rs, broker.rs, handler.rs, http/routes.rs}, and safekeeper/spec/*.tla;
    • storage_broker/proto/broker.proto;
    • pageserver/src/tenant/timeline/walreceiver/{connection_manager.rs, walreceiver_connection.rs}, pageserver/src/walingest.rs;
    • libs/wal_decoder/, libs/postgres_ffi/build.rs, libs/compute_api/src/{spec.rs, responses.rs};
    • compute_tools/src/{spec.rs, config.rs, http/server.rs, compute.rs};
    • storage_controller/src/{compute_hook.rs, http.rs};
    • proxy/src/control_plane/client/cplane_proxy_v1.rs, messages.rs;
    • docs/{walservice.md, core_changes.md, safekeeper-protocol.md} (partly out of date), and RFCs 009, 013, 025, 035 and 041;
    • vendor/revisions.json, .gitmodules.
  • Neon activity and statements:
    • the neondatabase/neon and neondatabase/postgres commit, branch and compare APIs;
    • Discussion #12835;
    • neon.com/docs/introduction/architecture-overview, neon.com/blog/paxos, neon.com/blog/wal-s3-lakebase-storage-for-the-era-of-agents;
    • jack-vanlightly.com (2023-11-15, 2025-02-19; no latency figures);
    • clickhouse.com/blog/postgresbench-ha.
  • CloudNativePG:
    • cloudnative-pg/cloudnative-pg (license, releases v1.30.1 and v1.30.0);
    • cloudnative-pg.io/docs/devel/{supported_releases, backup, recovery};
    • cloudnative-pg/plugin-barman-cloud (releases, docs/usage, docs/object_stores);
    • cloudnative-pg/postgres-containers;
    • cncf.io/projects/cloudnativepg.
  • PgDog: pgdogdev/pgdog (LICENSE, README.md, CONTRIBUTING.md, releases, CHANGELOG-ENTERPRISE.md); docs.pgdog.dev (databases, authentication, prepared statements, enterprise edition); pgdogdev/helm.
  • PgBouncer: pgbouncer/pgbouncer (COPYRIGHT, the 1.26.0 release).
  • Postgres: postgresql.org/support/versioning, /developer/roadmap, and the 2026-08-13 and 2026-05-14 release announcements.
  • Object storage and log latency:
    • AWS S3 performance guidelines ("tens of milliseconds");
    • the S3 Express One Zone product page ("single-digit millisecond");
    • docs.warpstream.com low-latency clusters, and warpstream.com/blog (2026-02-04);
    • docs.automq.com WAL storage.
  • TiKV:
    • pingcap/docs release-8.5: tikv-configuration-file.md, pd-configuration-file.md, configure-placement-rules.md, configure-load-base-split.md, latency-breakdown.md, tune-tikv-thread-performance.md, and the 5.0, 5.3 and 8.0 release notes;
    • tikv/tikv release-8.5 (src/storage/mod.rs, src/storage/txn/commands/atomic_store.rs, components/api_version/src/api_v2.rs), and tikv#10540;
    • tikv.org/docs/7.1/develop/rawkv/cas, tikv.org/docs/6.1/deploy/performance/overview;
    • the PingCAP Raft Engine blog;
    • tikv/client-rust (src/store/request.rs, PR #562), and tikv/client-go config/client.go;
    • bitsand.cloud/posts/cross-az-latencies; the AWS fault-isolation whitepaper (AZs); smalldatum.blogspot.com (2026-01, SSD power-loss protection and fsync); dina-kar/client-rust at 1f8962b (src/raw/client.rs, src/transaction/transaction.rs, src/config.rs); docs/plans/r1-dependency-spike.md §(i).
  • Loams:
    • §02 (WAL classes), §04, §20 (D126, D130), §22 §6.4, §23;
    • crates/loams-log/src/writer.rs (the standard write path: one flush in flight, ack after the PUT and CommitWal);
    • crates/loams-meta-tikv/src/leases.rs (check_fence);
    • crates/loams-tikv;
    • deploy/neon/, deploy/tikv/;
    • deny.toml;
    • D11, D60, D61, D72, D96, D126, D130, D148–D157, D-SC-6, D-SC-12, D-SC-16; Q30, Q45–Q49.

On this page