Implementation:Vllm project Vllm LLMEngine Add Request
| Knowledge Sources | |
|---|---|
| Domains | LLM Serving, Model Adaptation, Request Scheduling |
| Last Updated | 2026-02-08 13:00 GMT |
Overview
Concrete tool for submitting LoRA-augmented inference requests and executing batched inference provided by vllm.
Description
The LLMEngine.add_request() method (defined at lines 217-284 of vllm/v1/engine/llm_engine.py) accepts an inference request along with an optional LoRARequest specifying which adapter to apply. It validates the request ID, processes the raw prompt through the input processor (tokenization and preparation), and enqueues the request in both the output processor and engine core.
The LLMEngine.step() method (defined at lines 286-320) executes one iteration of the engine loop: it retrieves outputs from the engine core, processes them through the output processor, handles any stop-string-triggered aborts, records statistics, and returns a list of RequestOutput objects. Together, these two methods form the core request-processing loop for multi-LoRA serving.
When n > 1 is specified in sampling parameters, add_request fans out the request into n child requests via a ParentRequest wrapper, each independently scheduled and processed.
Usage
Use add_request() to submit individual inference requests with per-request LoRA adapters to the engine. Use step() in a loop to drive execution and collect outputs. The typical pattern is to submit all requests via add_request() and then loop on step() until has_unfinished_requests() returns False.
Code Reference
Source Location
- Repository: vllm
- File: vllm/v1/engine/llm_engine.py (lines 217-320)
Signature
def add_request(
self,
request_id: str,
prompt: EngineCoreRequest | PromptType | DictPrompt | TokPrompt,
params: SamplingParams | PoolingParams,
arrival_time: float | None = None,
lora_request: LoRARequest | None = None,
tokenization_kwargs: dict[str, Any] | None = None,
trace_headers: Mapping[str, str] | None = None,
priority: int = 0,
prompt_text: str | None = None,
) -> None
def step(self) -> list[RequestOutput | PoolingRequestOutput]
Import
from vllm.v1.engine.llm_engine import LLMEngine
I/O Contract
Inputs (add_request)
| Name | Type | Required | Description |
|---|---|---|---|
| request_id | str | Yes | Unique string identifier for this request. Must be unique across all active requests. |
| prompt | PromptType | Yes | The input prompt as a string, token list, or pre-processed EngineCoreRequest. |
| params | SamplingParams or PoolingParams | Yes | Sampling parameters controlling generation (temperature, top_k, max_tokens, etc.). |
| arrival_time | float or None | No | Timestamp of request arrival for scheduling priority. Default: None (uses current time). |
| lora_request | LoRARequest or None | No | LoRA adapter to apply for this request. None uses the base model. Default: None. |
| tokenization_kwargs | dict or None | No | Additional keyword arguments for the tokenizer. Default: None. |
| trace_headers | Mapping[str, str] or None | No | Trace headers for observability. Default: None. |
| priority | int | No | Request priority for scheduling. Lower values indicate higher priority. Default: 0. |
Outputs (step)
| Name | Type | Description |
|---|---|---|
| request_outputs | list[RequestOutput or PoolingRequestOutput] | List of outputs for requests that produced new tokens or finished during this step |
Usage Examples
Multi-LoRA Request Submission and Processing Loop
from vllm import EngineArgs, LLMEngine, SamplingParams, RequestOutput
from vllm.lora.request import LoRARequest
# Assume engine is already initialized with enable_lora=True
# and lora_path points to downloaded adapter weights
# Submit base model request (no adapter)
engine.add_request(
"req-0",
"A robot may not injure a human being",
SamplingParams(temperature=0.0, max_tokens=128),
lora_request=None,
)
# Submit request with first LoRA adapter
engine.add_request(
"req-1",
"[user] Write a SQL query... [/user] [assistant]",
SamplingParams(temperature=0.0, max_tokens=128),
lora_request=LoRARequest("sql-lora", 1, lora_path),
)
# Submit request with second LoRA adapter
engine.add_request(
"req-2",
"[user] Write a SQL query... [/user] [assistant]",
SamplingParams(temperature=0.0, max_tokens=128),
lora_request=LoRARequest("sql-lora2", 2, lora_path),
)
# Process all requests in a continuous batching loop
while engine.has_unfinished_requests():
request_outputs: list[RequestOutput] = engine.step()
for output in request_outputs:
if output.finished:
print(f"Request {output.request_id}: {output.outputs[0].text}")
Interleaved Submission and Processing
from vllm import LLMEngine, SamplingParams, RequestOutput
from vllm.lora.request import LoRARequest
# Submit requests one at a time, interleaved with step() calls
test_prompts = [
("Prompt 1", SamplingParams(temperature=0.0, max_tokens=128), None),
("Prompt 2", SamplingParams(temperature=0.8, max_tokens=128),
LoRARequest("sql-lora", 1, lora_path)),
]
request_id = 0
while test_prompts or engine.has_unfinished_requests():
if test_prompts:
prompt, params, lora_req = test_prompts.pop(0)
engine.add_request(str(request_id), prompt, params, lora_request=lora_req)
request_id += 1
request_outputs = engine.step()
for output in request_outputs:
if output.finished:
print(output)