On this page
Full text in Loam is two open-source projects working together. Tantivy is the search library: it builds and queries inverted indexes. Quickwit is a search engine built on Tantivy for object storage, and its split format is what lets Loam serve text from a bucket. Loam uses Tantivy as a library and vendors a set of Quickwit's modules; it does not run Quickwit.
| Repositories | quickwit-oss/tantivy, quickwit-oss/quickwit |
| Licenses | Tantivy: MIT. Quickwit: Apache-2.0, since it was relicensed from AGPL-3.0 in January 2025, after Datadog's acquisition |
| Versions | Tantivy 0.26.2, pinned exactly with its quickwit feature; Quickwit code vendored file by file from a pinned commit |
| Crates in Loam | operon-text (Loam's Tantivy integration), operon-quickwit (the vendored Quickwit code) |
| In Loam | Available |
How Tantivy works
Tantivy is a Rust library in the mould of Lucene.
- Segments. An index is a set of immutable segments. Each holds an inverted index, per-document stored data and columnar values. New documents go into new segments; merges combine small segments into larger ones and drop deleted documents.
- The inverted index. A term dictionary maps every term to its postings: the ids of the documents that contain it, with term frequencies and optionally positions for phrase queries. Postings are stored in compressed, bit-packed blocks with skip information.
- Fast fields are column-oriented values per document (numbers, dates, keywords), used for sorting, filtering and aggregations, like Lucene's doc values.
- Fieldnorms record each field's length per document, for BM25's length normalization.
- Scoring and top-k. Queries score with BM25. Top-k retrieval uses block-max WAND: each postings block records the maximum score any of its documents can contribute, so whole blocks that cannot reach the current top-k are skipped.
- Analyzers turn text into terms: tokenizers, lowercasing, stemming and stop words, configurable per field.
Quickwit's splits
Quickwit indexes logs and traces on object storage. Its unit of storage is a split: one immutable Tantivy index (typically a single segment) bundled with all its files into one object, with a small hotcache in its footer. The hotcache holds what a searcher needs before it can read anything else: the skeleton of the term dictionary, fast-field metadata and the offsets of every file inside the bundle.
So opening a split costs one ranged GET of the footer. After that, a query reads only the byte ranges it needs: the postings of its terms, some positions, some fast-field pages. Nothing is downloaded whole. That is what makes an inverted index on S3 practical.
Why Loam built on these, and how
Why Tantivy. It is the most mature full-text library in Rust, fast, actively maintained and MIT-licensed. ParadeDB's pg_search is built on a fork of it, but ParadeDB itself is AGPL, so we use Tantivy directly.
Why not Lance's full-text index. Loam needs Elasticsearch-compatible behaviour for its Elasticsearch surface: Lucene's standard and English analyzers token for token, BM25 statistics that match, aggregations over fast fields and highlighting. Tantivy gives us all of that, and Quickwit had already solved serving it from object storage; Lance's full-text index was not designed around Elasticsearch compatibility.
Why vendor Quickwit instead of running it. Quickwit is an append-only engine with its own metastore, ingest pipeline and control plane. Loam has its own metastore and log, and needs upserts and deletes, which Quickwit does not do. And Quickwit's internal crates have no stable API and follow Datadog's observability roadmap. So we took specific modules, file by file, from a pinned commit into operon-quickwit:
storageanddirectories: the split bundle, the hotcache, async and caching directories, warm-up;- the Elasticsearch Query DSL → Quickwit's query AST → Tantivy query path, and the doc-mapper query builder;
- aggregation merge glue, for combining aggregations across splits;
StableLogMergePolicy, Quickwit's log-structured merge policy;quickwit-datetime, for Elasticsearch-style date formats.
The vendored code is adapted to crates.io Tantivy, with Quickwit's license and notices kept.
How Loam uses them
- One split per indexing batch. A worker that applies a batch of writes to a collection writes one Lance fragment and one Tantivy split from the same records, and commits both under one collection manifest. The manifest's
SplitRefrecords the footer's byte range, so opening a split is exactly one GET. - Row ids connect the two formats. Each Tantivy document stores its primary key and its Lance stable row id. Documents are added in row-id order, so row id → (split, doc id) is a binary search.
- Deletes and upserts. Tantivy segments are immutable, so Loam keeps one roaring delete bitmap per split in its own small format, rewritten whole on each change. The primary-key index says which row, and therefore which split and document, to mark.
- Merges re-index. Split merges follow the vendored log-structured policy, bounded by split size and grouped by schema version. A merged split is rebuilt from the documents'
_source, so deleted documents disappear and fields are re-analyzed with the right schema. A split with 30% or more deleted documents is rewritten on its own. - Storage through Loam. Quickwit's
Storagetrait is implemented over Loam's object-store layer and its RAM and NVMe range cache, so split reads share the hot tier. Hot collections can pin their splits in the cache. - The tail. Records not yet in a split are indexed in an in-memory Tantivy index on the query node, so strong reads see them.
- Sparse vectors ride along. For the Qdrant API's sparse vectors, each split carries a postings field with one term per sparse index and a fast field holding the vector, and Loam scores candidates exactly.
Scores that do not depend on layout
Lucene, and therefore Elasticsearch, computes BM25 statistics per shard and counts deleted documents until a merge removes them. So a document's score changes when its neighbours are deleted or merged. Loam computes BM25 statistics globally and live-only: over every split in the manifest plus the tail, the document count excludes deleted and shadowed documents, document frequencies subtract deleted postings, and token counts use each live document's fieldnorm bucket. WAND prunes with a small slack, and every candidate is rescored with a canonical form of the query. The result: scores are bit-identical before and after a merge, and on a hot node or a cold one.
Limits
- Splits built with the
quickwitfeature are not readable by a Tantivy built without it, because the feature switches the term dictionary to an SSTable format. The workspace pins both. - Delete bitmaps are rewritten whole. Fine for typical splits (1 to 5 GiB targets); a collection with very frequent deletes pays for it until merges catch up.
- A vendored fork is ours to maintain. Upstream fixes are pulled in by hand.
- Elasticsearch language analyzers beyond standard and English (the Snowball stemmers and stop-word lists) are planned, each checked token for token against Elasticsearch's
_analyze.
The next post leaves collections for tables: Apache Iceberg, iceberg-rust and Lakekeeper.