Jump to content

Connect SuperML | Leeroopedia MCP: Equip your AI agents with best practices, code verification, and debugging knowledge. Powered by Leeroo — building Organizational Superintelligence. Contact us at founders@leeroo.com.

Principle:Vllm project Vllm Multi LoRA Request Routing

From Leeroopedia


Knowledge Sources
Domains LLM Serving, Model Adaptation, Request Scheduling
Last Updated 2026-02-08 13:00 GMT

Overview

Multi-LoRA request routing is the mechanism by which an inference engine accepts requests that each specify a different LoRA adapter and schedules them for execution with the correct adapter weights applied.

Description

In a multi-LoRA serving system, each inference request can optionally specify a LoRA adapter to be applied during its forward pass. The engine must accept these requests, associate each one with the correct adapter, and schedule execution such that the appropriate adapter weights are loaded into GPU memory when the request is processed.

The routing mechanism operates at two levels: the add_request method accepts a request along with its optional LoRA adapter specification and enqueues it for processing, while the step method executes a batch of requests and returns their outputs. The engine's scheduler is responsible for batching requests in a way that respects the max_loras constraint -- only requests whose adapters are currently loaded in GPU adapter slots can be processed in the same batch.

Requests without a LoRA adapter (where lora_request is None) use the base model weights and can be freely mixed with adapter-augmented requests in the same batch without consuming an adapter slot.

Usage

Use multi-LoRA request routing when:

  • Submitting inference requests that target different LoRA adapters to a single engine instance
  • Mixing base-model requests with adapter-augmented requests in a single serving loop
  • Building a continuous batching loop that processes requests with heterogeneous adapter requirements
  • Implementing an online serving pipeline where different users or tasks require different fine-tuned behaviors

Theoretical Basis

The request routing mechanism implements a producer-consumer pattern with adapter-aware scheduling:

Request Submission (add_request): Each call to add_request enqueues a single request with its associated prompt, sampling parameters, and optional LoRA adapter. The engine processes the raw inputs through its input processor, which tokenizes the prompt and prepares it for execution. The request is then added to both the output processor (for tracking) and the engine core (for scheduling and execution).

Batch Execution (step): Each call to step triggers the engine core to select a batch of pending requests, execute a forward pass, and return incremental outputs. The scheduler selects requests based on available memory, sequence lengths, and adapter slot availability. If a request's adapter is not currently loaded in a GPU slot, it may be deferred to a future step while the adapter is swapped in.

Fan-Out for n>1: When a request specifies n > 1 (multiple completions), the engine fans out the request into n child requests that share a parent, each processed independently with its own sampling parameters.

Continuous Batching: The step method is designed to be called in a loop. Each invocation returns a list of RequestOutput objects for requests that produced new tokens or finished. The loop continues until all requests are complete, as indicated by engine.has_unfinished_requests().

Related Pages

Implemented By

Page Connections

Double-click a node to navigate. Hold to expand connections.
Principle
Implementation
Heuristic
Environment