Implementation:Lance format Lance MemtableReadBench
| Knowledge Sources | |
|---|---|
| Domains | Benchmarking, Performance |
| Last Updated | 2026-02-08 19:33 GMT |
Overview
Description
MemtableReadBench is a Criterion-based benchmark that compares read performance between in-memory MemTables (using MemTableScanner) and in-memory Lance tables. It exercises four distinct read operations, each comparing the MemTable path against a conventional Lance dataset:
- Scan — Full table scan returning all rows, measuring total throughput.
- Point Lookup — Scalar index-based point lookups using a BTree index on the primary key column.
- Full-Text Search (FTS) — Token-based text search using an inverted index on a text column.
- Vector Search — IVF-PQ vector similarity search, comparing a MemTable with trained IVF centroids and PQ codebook against a Lance dataset with a materialized IVF-PQ index.
The data schema includes three columns:
id(Int64) — Primary keytext(Utf8) — Text column with common word patterns for FTS testingvector(FixedSizeList<Float32>) — Normalized random vectors
The benchmark trains IVF centroids and a PQ codebook from the generated data to build MemTable-level vector search indexes. For Lance dataset comparisons, it creates equivalent materialized indexes.
Usage
This benchmark is used to evaluate whether the MemTable in-memory data structure can match or exceed the read performance of a fully materialized Lance table for different access patterns. It helps developers identify MemTable performance bottlenecks and optimize the MemTable, IndexStore, and CacheConfig components.
Code Reference
Source Location
rust/lance/benches/memtable_read.rs (1119 lines)
Signature
The benchmark defines a single orchestrator that delegates to four sub-benchmarks:
fn all_benchmarks(c: &mut Criterion) {
bench_scan(c);
bench_point_lookup(c);
bench_fts(c);
bench_vector_search(c);
}
Import
use lance::dataset::mem_wal::write::{CacheConfig, IndexStore, MemTable};
use lance::dataset::{Dataset, WriteParams};
use lance::index::vector::VectorIndexParams;
use lance_index::scalar::inverted::tokenizer::InvertedIndexParams;
use lance_index::scalar::FullTextSearchQuery;
use lance_index::vector::ivf::storage::IvfModel;
use lance_index::vector::ivf::IvfBuildParams;
use lance_index::vector::kmeans::{train_kmeans, KMeansParams};
use lance_index::vector::pq::builder::PQBuildParams;
use lance_index::{DatasetIndexExt, IndexType};
use criterion::{criterion_group, criterion_main, BenchmarkId, Criterion, Throughput};
I/O Contract
Inputs
| Parameter | Type | Default | Description |
|---|---|---|---|
NUM_ROWS |
Environment variable | 10000 | Total number of rows in the dataset |
BATCH_SIZE |
Environment variable | 100 | Number of rows per batch |
VECTOR_DIM |
Environment variable | 128 | Dimension of the vector column |
SAMPLE_SIZE |
Environment variable | 100 | Number of benchmark iterations (minimum 10) |
Outputs
| Output | Type | Description |
|---|---|---|
| Criterion report | HTML/JSON | Throughput statistics (rows/sec) for each benchmark group comparing MemTable vs Lance, with optional flamegraph profiling on Linux |
| Comparison data | Criterion stats | Side-by-side MemTable vs Lance dataset performance for scan, point lookup, FTS, and vector search |
Usage Examples
Run the full benchmark suite:
cargo bench -p lance --bench memtable_read
Run with custom data size:
export NUM_ROWS=50000
export VECTOR_DIM=256
cargo bench -p lance --bench memtable_read
Run only the vector search sub-benchmark:
cargo bench -p lance --bench memtable_read -- "vector"
Run only the FTS sub-benchmark:
cargo bench -p lance --bench memtable_read -- "fts"
Related Pages
- Lance_format_Lance_MemWalReadBench — Benchmark for LSM scanner read performance with MemWAL
- Lance_format_Lance_MemWalWriteBench — Benchmark for MemWAL write throughput