Jump to content

Connect SuperML | Leeroopedia MCP: Equip your AI agents with best practices, code verification, and debugging knowledge. Powered by Leeroo — building Organizational Superintelligence. Contact us at founders@leeroo.com.

Implementation:LMCache LMCache VLLM Multi Process Adapter

From Leeroopedia
Revision as of 15:26, 16 February 2026 by Admin (talk | contribs) (Auto-imported from implementations/LMCache_LMCache_VLLM_Multi_Process_Adapter.md)
(diff) ← Older revision | Latest revision (diff) | Newer revision → (diff)


Knowledge Sources
Domains vLLM Integration, Multi-Process Architecture
Last Updated 2026-02-09 00:00 GMT

Overview

This module provides multi-process adapter classes that enable vLLM's scheduler and GPU workers to communicate with a separate LMCache server process via ZMQ message queues for KV cache lookup, store, and retrieve operations.

Description

The vllm_multi_process_adapter.py module implements two adapter classes -- LMCacheMPSchedulerAdapter for the vLLM scheduler process and LMCacheMPWorkerAdapter for GPU worker processes. These adapters communicate with a dedicated LMCache multiprocess server over ZMQ using MessageQueueClient. The scheduler adapter handles asynchronous lookup requests and prefix matching, while the worker adapter manages KV cache registration, batched store/retrieve operations with CUDA event synchronization, and tracking of finished requests across the vLLM-LMCache boundary.

Usage

Use LMCacheMPSchedulerAdapter in the vLLM scheduler to submit lookup requests and check prefix match results. Use LMCacheMPWorkerAdapter in GPU workers to register KV caches and submit store/retrieve operations. Both require a running LMCache multiprocess server.

Code Reference

Source Location

Signature

def wrap_kv_caches(kv_caches: dict[str, torch.Tensor]) -> KVCache: ...

def send_lmcache_request(
    mq_client: MessageQueueClient, request_type: RequestType,
    payloads: list[Any],
) -> MessagingFuture[Any]: ...

def get_lmcache_chunk_size(mq_client: MessageQueueClient) -> int: ...

def striding_block_hashes(
    block_hashes: list[bytes], blocks_in_chunk: int,
) -> Iterable[bytes]: ...

@dataclass
class LoadStoreOp:
    block_hashes: list[bytes]
    block_ids: list[int]

class LMCacheMPSchedulerAdapter:
    def __init__(
        self, server_url: str, context: zmq.Context,
        model_name: str, world_size: int, kv_rank: int,
        vllm_block_size: int,
    ): ...
    def maybe_submit_lookup_request(
        self, request_id: str, block_hashes: list[bytes]
    ): ...
    def check_lookup_result(self, request_id: str) -> int | None: ...
    def num_blocks_per_chunk(self) -> int: ...
    def cleanup_lookup_result(self, request_id: str) -> None: ...

class LMCacheMPWorkerAdapter:
    def __init__(
        self, server_url: str, context: zmq.Context,
        model_name: str, world_size: int, kv_rank: int,
        vllm_block_size: int,
    ): ...
    def register_kv_caches(
        self, kv_caches: dict[str, torch.Tensor]
    ): ...
    def submit_store_request(
        self, request_id: str, op: LoadStoreOp,
        event: torch.cuda.Event,
    ): ...
    def submit_retrieve_request(
        self, request_id: str, op: LoadStoreOp,
        event: torch.cuda.Event,
    ): ...
    def batched_submit_store_requests(
        self, request_ids: list[str], ops: list[LoadStoreOp],
        event: torch.cuda.Event,
    ): ...
    def batched_submit_retrieve_requests(
        self, request_ids: list[str], ops: list[LoadStoreOp],
        event: torch.cuda.Event,
    ): ...
    def get_finished(
        self, finished_req_ids_from_engine: set[str]
    ) -> tuple[set[str] | None, set[str] | None]: ...
    def num_blocks_per_chunk(self) -> int: ...
    def shutdown(self): ...

Import

from lmcache.integration.vllm.vllm_multi_process_adapter import (
    LMCacheMPSchedulerAdapter,
    LMCacheMPWorkerAdapter,
    LoadStoreOp,
    wrap_kv_caches,
    send_lmcache_request,
    striding_block_hashes,
)

I/O Contract

Inputs

Name Type Required Description
server_url str Yes ZMQ server URL for the LMCache multiprocess server
context zmq.Context Yes ZMQ context for socket creation
model_name str Yes Model name for constructing LMCache keys
world_size int Yes World size for LMCache key construction
kv_rank int Yes KV rank (worker ID) for LMCache key construction
vllm_block_size int Yes Block size used in vLLM (must divide LMCache chunk size)
block_hashes list[bytes] Yes (for operations) Block hashes identifying cache chunks
event torch.cuda.Event Yes (for store/retrieve) CUDA event recorded after the model inference step

Outputs

Name Type Description
lookup_result int or None Total number of tokens matched via prefix matching, or None if not ready
finished_stores set[str] Set of request IDs whose store operations have completed
finished_retrieves set[str] Set of request IDs whose retrieve operations have completed
blocks_in_chunk int Number of vLLM blocks per LMCache chunk

Usage Examples

import zmq
from lmcache.integration.vllm.vllm_multi_process_adapter import (
    LMCacheMPSchedulerAdapter, LMCacheMPWorkerAdapter, LoadStoreOp,
)

# Scheduler side: lookup KV cache matches
ctx = zmq.Context()
scheduler_adapter = LMCacheMPSchedulerAdapter(
    server_url="tcp://localhost:5555",
    context=ctx,
    model_name="meta-llama/Llama-3.1-8B",
    world_size=1,
    kv_rank=0,
    vllm_block_size=16,
)

scheduler_adapter.maybe_submit_lookup_request("req-1", block_hashes)
result = scheduler_adapter.check_lookup_result("req-1")
if result is not None:
    print(f"Matched {result} tokens from cache")

# Worker side: register KV caches and store
worker_adapter = LMCacheMPWorkerAdapter(
    server_url="tcp://localhost:5555",
    context=ctx,
    model_name="meta-llama/Llama-3.1-8B",
    world_size=1,
    kv_rank=0,
    vllm_block_size=16,
)
worker_adapter.register_kv_caches(kv_cache_dict)

op = LoadStoreOp(block_hashes=hashes, block_ids=ids)
worker_adapter.submit_store_request("req-1", op, cuda_event)

Page Connections

Double-click a node to navigate. Hold to expand connections.
Principle
Implementation
Heuristic
Environment