Jump to content

Connect SuperML | Leeroopedia MCP: Equip your AI agents with best practices, code verification, and debugging knowledge. Powered by Leeroo — building Organizational Superintelligence. Contact us at founders@leeroo.com.

Implementation:Lance format Lance MemWalReadBench

From Leeroopedia


Knowledge Sources
Domains Benchmarking, Performance
Last Updated 2026-02-08 19:33 GMT

Overview

Description

MemWalReadBench is a Criterion-based benchmark that measures LSM (Log-Structured Merge) scanner read performance in the Lance MemWAL subsystem. It compares scanning performance between a single base Lance table (baseline) and an LSM scan that combines a base table with flushed MemTables and an active MemTable.

The benchmark sets up a BenchContext containing:

  • A base dataset with configurable row count
  • Two flushed MemTable generations
  • One active MemTable

Four benchmark groups exercise different read patterns:

  • LSM Scan — Full table scan with and without memtables, measuring total row throughput.
  • LSM Scan Projected — Scan with column projection (selecting a subset of columns).
  • LSM Point Lookup — Primary key-based point lookups using the LsmPointLookupPlanner.
  • LSM Vector Search — KNN vector similarity search across LSM levels using LsmVectorSearchPlanner.

Usage

This benchmark is used to evaluate the read overhead introduced by the MemWAL layer compared to direct table access. It helps developers understand the performance cost of layered LSM scanning and to optimize the LsmScanner, LsmPointLookupPlanner, and LsmVectorSearchPlanner components.

Code Reference

Source Location

rust/lance/benches/mem_wal_read.rs (1059 lines)

Signature

The benchmark defines a single orchestrator function that delegates to four sub-benchmarks:

fn all_benchmarks(c: &mut Criterion) {
    bench_scan(c);
    bench_scan_with_projection(c);
    bench_point_lookup(c);
    bench_vector_search(c);
}

Import

use lance::dataset::mem_wal::scanner::{
    ActiveMemTableRef, LsmDataSourceCollector, LsmPointLookupPlanner,
    LsmScanner, LsmVectorSearchPlanner, RegionSnapshot,
};
use lance::dataset::mem_wal::{DatasetMemWalExt, MemWalConfig, RegionWriterConfig};
use lance::dataset::{Dataset, WriteParams};
use lance_linalg::distance::DistanceType;
use criterion::{criterion_group, criterion_main, BenchmarkId, Criterion, Throughput};

I/O Contract

Inputs

Parameter Type Default Description
DATASET_PREFIX Environment variable Temp directory Base URI for datasets (supports S3, GCS, Azure, or local path)
BASE_ROWS Environment variable 10000 Number of rows in the base table
MEMTABLE_ROWS Environment variable 1000 Number of rows per MemTable generation
BATCH_SIZE Environment variable 100 Rows per write batch
SAMPLE_SIZE Environment variable 100 Number of benchmark iterations (minimum 10)
VECTOR_DIM Environment variable 128 Vector dimension for the vector search benchmark

Outputs

Output Type Description
Criterion report HTML/JSON Throughput statistics (rows/sec) for each benchmark group, with optional flamegraph profiling on Linux
Comparison data Criterion stats Compares "base_only" scan vs "lsm" scan with memtable layers

Usage Examples

Run the full benchmark suite against a local temporary directory:

cargo bench -p lance --bench mem_wal_read

Run against S3 storage:

export AWS_DEFAULT_REGION=us-east-1
export DATASET_PREFIX=s3://your-bucket/bench/mem_wal_read
cargo bench -p lance --bench mem_wal_read

Run against a specific local directory with custom parameters:

export DATASET_PREFIX=/tmp/bench/mem_wal_read
export BASE_ROWS=50000
export MEMTABLE_ROWS=5000
cargo bench -p lance --bench mem_wal_read

Run only the vector search sub-benchmark:

cargo bench -p lance --bench mem_wal_read -- "vector"

Related Pages

Page Connections

Double-click a node to navigate. Hold to expand connections.
Principle
Implementation
Heuristic
Environment