Implementation:Lance format Lance MemWalReadBench
| Knowledge Sources | |
|---|---|
| Domains | Benchmarking, Performance |
| Last Updated | 2026-02-08 19:33 GMT |
Overview
Description
MemWalReadBench is a Criterion-based benchmark that measures LSM (Log-Structured Merge) scanner read performance in the Lance MemWAL subsystem. It compares scanning performance between a single base Lance table (baseline) and an LSM scan that combines a base table with flushed MemTables and an active MemTable.
The benchmark sets up a BenchContext containing:
- A base dataset with configurable row count
- Two flushed MemTable generations
- One active MemTable
Four benchmark groups exercise different read patterns:
- LSM Scan — Full table scan with and without memtables, measuring total row throughput.
- LSM Scan Projected — Scan with column projection (selecting a subset of columns).
- LSM Point Lookup — Primary key-based point lookups using the
LsmPointLookupPlanner. - LSM Vector Search — KNN vector similarity search across LSM levels using
LsmVectorSearchPlanner.
Usage
This benchmark is used to evaluate the read overhead introduced by the MemWAL layer compared to direct table access. It helps developers understand the performance cost of layered LSM scanning and to optimize the LsmScanner, LsmPointLookupPlanner, and LsmVectorSearchPlanner components.
Code Reference
Source Location
rust/lance/benches/mem_wal_read.rs (1059 lines)
Signature
The benchmark defines a single orchestrator function that delegates to four sub-benchmarks:
fn all_benchmarks(c: &mut Criterion) {
bench_scan(c);
bench_scan_with_projection(c);
bench_point_lookup(c);
bench_vector_search(c);
}
Import
use lance::dataset::mem_wal::scanner::{
ActiveMemTableRef, LsmDataSourceCollector, LsmPointLookupPlanner,
LsmScanner, LsmVectorSearchPlanner, RegionSnapshot,
};
use lance::dataset::mem_wal::{DatasetMemWalExt, MemWalConfig, RegionWriterConfig};
use lance::dataset::{Dataset, WriteParams};
use lance_linalg::distance::DistanceType;
use criterion::{criterion_group, criterion_main, BenchmarkId, Criterion, Throughput};
I/O Contract
Inputs
| Parameter | Type | Default | Description |
|---|---|---|---|
DATASET_PREFIX |
Environment variable | Temp directory | Base URI for datasets (supports S3, GCS, Azure, or local path) |
BASE_ROWS |
Environment variable | 10000 | Number of rows in the base table |
MEMTABLE_ROWS |
Environment variable | 1000 | Number of rows per MemTable generation |
BATCH_SIZE |
Environment variable | 100 | Rows per write batch |
SAMPLE_SIZE |
Environment variable | 100 | Number of benchmark iterations (minimum 10) |
VECTOR_DIM |
Environment variable | 128 | Vector dimension for the vector search benchmark |
Outputs
| Output | Type | Description |
|---|---|---|
| Criterion report | HTML/JSON | Throughput statistics (rows/sec) for each benchmark group, with optional flamegraph profiling on Linux |
| Comparison data | Criterion stats | Compares "base_only" scan vs "lsm" scan with memtable layers |
Usage Examples
Run the full benchmark suite against a local temporary directory:
cargo bench -p lance --bench mem_wal_read
Run against S3 storage:
export AWS_DEFAULT_REGION=us-east-1
export DATASET_PREFIX=s3://your-bucket/bench/mem_wal_read
cargo bench -p lance --bench mem_wal_read
Run against a specific local directory with custom parameters:
export DATASET_PREFIX=/tmp/bench/mem_wal_read
export BASE_ROWS=50000
export MEMTABLE_ROWS=5000
cargo bench -p lance --bench mem_wal_read
Run only the vector search sub-benchmark:
cargo bench -p lance --bench mem_wal_read -- "vector"
Related Pages
- Lance_format_Lance_MemWalWriteBench — Benchmark for MemWAL write throughput
- Lance_format_Lance_MemtableReadBench — Benchmark comparing MemTable vs in-memory Lance table read performance