Implementation:InternLM Lmdeploy Profile Pipeline Api
| Knowledge Sources | |
|---|---|
| Domains | Benchmarking, Performance, Pipeline |
| Last Updated | 2026-02-07 15:00 GMT |
Overview
A benchmarking script that profiles the throughput and latency of lmdeploy's Python pipeline API by sending batched requests through the local inference engine.
Description
The profile_pipeline_api.py script measures the performance of lmdeploy's pipeline interface for local inference (no server required). It supports both TurboMind and PyTorch backends and provides two dataset sampling strategies:
- ShareGPT sampling (
sample_sharegpt_requests): Loads conversations from the ShareGPT dataset, filters by token length (prompt <= 1024, total <= 2048), and uses actual conversation lengths or a fixed output length. - Random sampling (
sample_random_requests): Generates requests with specified input/output token lengths using ShareGPT prompts as seed text, repeating/truncating to match desired lengths.
The Engine class wraps the lmdeploy.pipeline interface and processes all requests in a single batch. It supports both streaming (stream_infer) and non-streaming inference modes. Each request is tracked via a Session object from the Profiler class, recording token-level timing for metrics computation.
The script uses the Profiler class to compute and report:
- End-to-end latency (mean, P50, P75, P95, P99)
- Time to first token (TTFT)
- Time per output token (TPOT)
- Inter-token latency (ITL)
- Input/output throughput (tokens/second)
- Request throughput (requests/second)
Results are optionally saved to CSV with configurable hyperparameters.
Usage
Used for benchmarking lmdeploy pipeline API performance. Run from the command line with a dataset path and model path.
Code Reference
Source Location
- Repository: InternLM_Lmdeploy
- File: benchmark/profile_pipeline_api.py
- Lines: 1-368
Signature
def sample_sharegpt_requests(
dataset_path: str, num_requests: int,
tokenizer: PreTrainedTokenizerBase,
fixed_output_len: Optional[int] = None,
) -> List[Tuple[str, int, int]]: ...
def sample_random_requests(
input_len: int, output_len: int, num_prompts: int,
range_ratio: float, tokenizer: PreTrainedTokenizerBase,
dataset_path: str,
) -> List[Tuple[str, int, int]]: ...
class Engine:
def __init__(self, model_path: str, engine_config, csv: str): ...
def process_request(self, requests, profiler: Profiler,
temperature, top_p, top_k, stream_output): ...
def main(): ...
Import
from lmdeploy import GenerationConfig, PytorchEngineConfig, TurbomindEngineConfig, pipeline
from lmdeploy.profiler import Profiler, Session
I/O Contract
Inputs
| Name | Type | Required | Description |
|---|---|---|---|
| dataset | str | Yes | Path to the ShareGPT JSON dataset file |
| model_path | str | Yes | Path to the model or HuggingFace repo ID |
| --concurrency / -c | int | No | Number of working threads (default: 256) |
| --num-prompts / -n | int | No | Number of prompts to process (default: 5000) |
| --backend | str | No | Backend: turbomind or pytorch (default: turbomind) |
| --dataset-name | str | No | Dataset type: sharegpt or random (default: sharegpt) |
| --stream-output | flag | No | Enable streaming output mode |
| --csv | str | No | Output CSV file path (default: ./profile_pipeline_api.csv) |
| --tp | int | No | Tensor parallelism degree |
| --cache-max-entry-count | float | No | Cache max entry count ratio |
| --model-format | str | No | Model format (default: hf) |
Outputs
| Name | Type | Description |
|---|---|---|
| Console summary | text | Formatted table with latency/throughput metrics |
| CSV file | file | Benchmark results with hyperparameters and metrics |
Usage Examples
# Profile with ShareGPT dataset using TurboMind backend
# python benchmark/profile_pipeline_api.py \
# /path/to/ShareGPT.json \
# internlm/internlm2_5-7b-chat \
# --backend turbomind \
# --tp 1 \
# --concurrency 256 \
# --num-prompts 1000 \
# --stream-output \
# --csv results.csv
# Profile with random dataset
# python benchmark/profile_pipeline_api.py \
# /path/to/ShareGPT.json \
# internlm/internlm2_5-7b-chat \
# --backend pytorch \
# --dataset-name random \
# --random-input-len 512 \
# --random-output-len 256 \
# --num-prompts 500