Implementation:LMCache LMCache VLLM Multi Process Adapter
| Knowledge Sources | |
|---|---|
| Domains | vLLM Integration, Multi-Process Architecture |
| Last Updated | 2026-02-09 00:00 GMT |
Overview
This module provides multi-process adapter classes that enable vLLM's scheduler and GPU workers to communicate with a separate LMCache server process via ZMQ message queues for KV cache lookup, store, and retrieve operations.
Description
The vllm_multi_process_adapter.py module implements two adapter classes -- LMCacheMPSchedulerAdapter for the vLLM scheduler process and LMCacheMPWorkerAdapter for GPU worker processes. These adapters communicate with a dedicated LMCache multiprocess server over ZMQ using MessageQueueClient. The scheduler adapter handles asynchronous lookup requests and prefix matching, while the worker adapter manages KV cache registration, batched store/retrieve operations with CUDA event synchronization, and tracking of finished requests across the vLLM-LMCache boundary.
Usage
Use LMCacheMPSchedulerAdapter in the vLLM scheduler to submit lookup requests and check prefix match results. Use LMCacheMPWorkerAdapter in GPU workers to register KV caches and submit store/retrieve operations. Both require a running LMCache multiprocess server.
Code Reference
Source Location
- Repository: LMCache
- File: lmcache/integration/vllm/vllm_multi_process_adapter.py
- Lines: 1-519
Signature
def wrap_kv_caches(kv_caches: dict[str, torch.Tensor]) -> KVCache: ...
def send_lmcache_request(
mq_client: MessageQueueClient, request_type: RequestType,
payloads: list[Any],
) -> MessagingFuture[Any]: ...
def get_lmcache_chunk_size(mq_client: MessageQueueClient) -> int: ...
def striding_block_hashes(
block_hashes: list[bytes], blocks_in_chunk: int,
) -> Iterable[bytes]: ...
@dataclass
class LoadStoreOp:
block_hashes: list[bytes]
block_ids: list[int]
class LMCacheMPSchedulerAdapter:
def __init__(
self, server_url: str, context: zmq.Context,
model_name: str, world_size: int, kv_rank: int,
vllm_block_size: int,
): ...
def maybe_submit_lookup_request(
self, request_id: str, block_hashes: list[bytes]
): ...
def check_lookup_result(self, request_id: str) -> int | None: ...
def num_blocks_per_chunk(self) -> int: ...
def cleanup_lookup_result(self, request_id: str) -> None: ...
class LMCacheMPWorkerAdapter:
def __init__(
self, server_url: str, context: zmq.Context,
model_name: str, world_size: int, kv_rank: int,
vllm_block_size: int,
): ...
def register_kv_caches(
self, kv_caches: dict[str, torch.Tensor]
): ...
def submit_store_request(
self, request_id: str, op: LoadStoreOp,
event: torch.cuda.Event,
): ...
def submit_retrieve_request(
self, request_id: str, op: LoadStoreOp,
event: torch.cuda.Event,
): ...
def batched_submit_store_requests(
self, request_ids: list[str], ops: list[LoadStoreOp],
event: torch.cuda.Event,
): ...
def batched_submit_retrieve_requests(
self, request_ids: list[str], ops: list[LoadStoreOp],
event: torch.cuda.Event,
): ...
def get_finished(
self, finished_req_ids_from_engine: set[str]
) -> tuple[set[str] | None, set[str] | None]: ...
def num_blocks_per_chunk(self) -> int: ...
def shutdown(self): ...
Import
from lmcache.integration.vllm.vllm_multi_process_adapter import (
LMCacheMPSchedulerAdapter,
LMCacheMPWorkerAdapter,
LoadStoreOp,
wrap_kv_caches,
send_lmcache_request,
striding_block_hashes,
)
I/O Contract
Inputs
| Name | Type | Required | Description |
|---|---|---|---|
| server_url | str | Yes | ZMQ server URL for the LMCache multiprocess server |
| context | zmq.Context | Yes | ZMQ context for socket creation |
| model_name | str | Yes | Model name for constructing LMCache keys |
| world_size | int | Yes | World size for LMCache key construction |
| kv_rank | int | Yes | KV rank (worker ID) for LMCache key construction |
| vllm_block_size | int | Yes | Block size used in vLLM (must divide LMCache chunk size) |
| block_hashes | list[bytes] | Yes (for operations) | Block hashes identifying cache chunks |
| event | torch.cuda.Event | Yes (for store/retrieve) | CUDA event recorded after the model inference step |
Outputs
| Name | Type | Description |
|---|---|---|
| lookup_result | int or None | Total number of tokens matched via prefix matching, or None if not ready |
| finished_stores | set[str] | Set of request IDs whose store operations have completed |
| finished_retrieves | set[str] | Set of request IDs whose retrieve operations have completed |
| blocks_in_chunk | int | Number of vLLM blocks per LMCache chunk |
Usage Examples
import zmq
from lmcache.integration.vllm.vllm_multi_process_adapter import (
LMCacheMPSchedulerAdapter, LMCacheMPWorkerAdapter, LoadStoreOp,
)
# Scheduler side: lookup KV cache matches
ctx = zmq.Context()
scheduler_adapter = LMCacheMPSchedulerAdapter(
server_url="tcp://localhost:5555",
context=ctx,
model_name="meta-llama/Llama-3.1-8B",
world_size=1,
kv_rank=0,
vllm_block_size=16,
)
scheduler_adapter.maybe_submit_lookup_request("req-1", block_hashes)
result = scheduler_adapter.check_lookup_result("req-1")
if result is not None:
print(f"Matched {result} tokens from cache")
# Worker side: register KV caches and store
worker_adapter = LMCacheMPWorkerAdapter(
server_url="tcp://localhost:5555",
context=ctx,
model_name="meta-llama/Llama-3.1-8B",
world_size=1,
kv_rank=0,
vllm_block_size=16,
)
worker_adapter.register_kv_caches(kv_cache_dict)
op = LoadStoreOp(block_hashes=hashes, block_ids=ids)
worker_adapter.submit_store_request("req-1", op, cuda_event)