Jump to content

Connect SuperML | Leeroopedia MCP: Equip your AI agents with best practices, code verification, and debugging knowledge. Powered by Leeroo — building Organizational Superintelligence. Contact us at founders@leeroo.com.

Implementation:Vllm project Vllm LLMEngine Add Request

From Leeroopedia
Revision as of 17:05, 16 February 2026 by Admin (talk | contribs) (Auto-imported from implementations/Vllm_project_Vllm_LLMEngine_Add_Request.md)
(diff) ← Older revision | Latest revision (diff) | Newer revision → (diff)


Knowledge Sources
Domains LLM Serving, Model Adaptation, Request Scheduling
Last Updated 2026-02-08 13:00 GMT

Overview

Concrete tool for submitting LoRA-augmented inference requests and executing batched inference provided by vllm.

Description

The LLMEngine.add_request() method (defined at lines 217-284 of vllm/v1/engine/llm_engine.py) accepts an inference request along with an optional LoRARequest specifying which adapter to apply. It validates the request ID, processes the raw prompt through the input processor (tokenization and preparation), and enqueues the request in both the output processor and engine core.

The LLMEngine.step() method (defined at lines 286-320) executes one iteration of the engine loop: it retrieves outputs from the engine core, processes them through the output processor, handles any stop-string-triggered aborts, records statistics, and returns a list of RequestOutput objects. Together, these two methods form the core request-processing loop for multi-LoRA serving.

When n > 1 is specified in sampling parameters, add_request fans out the request into n child requests via a ParentRequest wrapper, each independently scheduled and processed.

Usage

Use add_request() to submit individual inference requests with per-request LoRA adapters to the engine. Use step() in a loop to drive execution and collect outputs. The typical pattern is to submit all requests via add_request() and then loop on step() until has_unfinished_requests() returns False.

Code Reference

Source Location

  • Repository: vllm
  • File: vllm/v1/engine/llm_engine.py (lines 217-320)

Signature

def add_request(
    self,
    request_id: str,
    prompt: EngineCoreRequest | PromptType | DictPrompt | TokPrompt,
    params: SamplingParams | PoolingParams,
    arrival_time: float | None = None,
    lora_request: LoRARequest | None = None,
    tokenization_kwargs: dict[str, Any] | None = None,
    trace_headers: Mapping[str, str] | None = None,
    priority: int = 0,
    prompt_text: str | None = None,
) -> None

def step(self) -> list[RequestOutput | PoolingRequestOutput]

Import

from vllm.v1.engine.llm_engine import LLMEngine

I/O Contract

Inputs (add_request)

Name Type Required Description
request_id str Yes Unique string identifier for this request. Must be unique across all active requests.
prompt PromptType Yes The input prompt as a string, token list, or pre-processed EngineCoreRequest.
params SamplingParams or PoolingParams Yes Sampling parameters controlling generation (temperature, top_k, max_tokens, etc.).
arrival_time float or None No Timestamp of request arrival for scheduling priority. Default: None (uses current time).
lora_request LoRARequest or None No LoRA adapter to apply for this request. None uses the base model. Default: None.
tokenization_kwargs dict or None No Additional keyword arguments for the tokenizer. Default: None.
trace_headers Mapping[str, str] or None No Trace headers for observability. Default: None.
priority int No Request priority for scheduling. Lower values indicate higher priority. Default: 0.

Outputs (step)

Name Type Description
request_outputs list[RequestOutput or PoolingRequestOutput] List of outputs for requests that produced new tokens or finished during this step

Usage Examples

Multi-LoRA Request Submission and Processing Loop

from vllm import EngineArgs, LLMEngine, SamplingParams, RequestOutput
from vllm.lora.request import LoRARequest

# Assume engine is already initialized with enable_lora=True
# and lora_path points to downloaded adapter weights

# Submit base model request (no adapter)
engine.add_request(
    "req-0",
    "A robot may not injure a human being",
    SamplingParams(temperature=0.0, max_tokens=128),
    lora_request=None,
)

# Submit request with first LoRA adapter
engine.add_request(
    "req-1",
    "[user] Write a SQL query... [/user] [assistant]",
    SamplingParams(temperature=0.0, max_tokens=128),
    lora_request=LoRARequest("sql-lora", 1, lora_path),
)

# Submit request with second LoRA adapter
engine.add_request(
    "req-2",
    "[user] Write a SQL query... [/user] [assistant]",
    SamplingParams(temperature=0.0, max_tokens=128),
    lora_request=LoRARequest("sql-lora2", 2, lora_path),
)

# Process all requests in a continuous batching loop
while engine.has_unfinished_requests():
    request_outputs: list[RequestOutput] = engine.step()
    for output in request_outputs:
        if output.finished:
            print(f"Request {output.request_id}: {output.outputs[0].text}")

Interleaved Submission and Processing

from vllm import LLMEngine, SamplingParams, RequestOutput
from vllm.lora.request import LoRARequest

# Submit requests one at a time, interleaved with step() calls
test_prompts = [
    ("Prompt 1", SamplingParams(temperature=0.0, max_tokens=128), None),
    ("Prompt 2", SamplingParams(temperature=0.8, max_tokens=128),
     LoRARequest("sql-lora", 1, lora_path)),
]

request_id = 0
while test_prompts or engine.has_unfinished_requests():
    if test_prompts:
        prompt, params, lora_req = test_prompts.pop(0)
        engine.add_request(str(request_id), prompt, params, lora_request=lora_req)
        request_id += 1

    request_outputs = engine.step()
    for output in request_outputs:
        if output.finished:
            print(output)

Related Pages

Implements Principle

Page Connections

Double-click a node to navigate. Hold to expand connections.
Principle
Implementation
Heuristic
Environment