Implementation:Turboderp org Exllamav2 Ext QMLP H
| Knowledge Sources | |
|---|---|
| Domains | MLP, Quantization, C_Extension |
| Last Updated | 2026-02-15 00:00 GMT |
Overview
C++ header declaring the API for ExLlamaV2's fused quantized MLP and Mixture-of-Experts (MoE) MLP modules, including construction, forward pass, LoRA support, and tensor-parallel variants.
Description
ext_qmlp.h defines the interface for two quantized feed-forward network subsystems:
make_q_mlp constructs a standard quantized MLP module. It accepts parameters including:
- layernorm and layernorm_bias -- Pre-MLP normalization weights.
- layernorm_is_rms -- Flag selecting RMS norm vs standard layer norm.
- q_gate, q_up, q_down -- Opaque handles to quantized weight matrices for the gate, up-projection, and down-projection layers.
- temp_state, temp_a, temp_b, temp_dq -- Scratch buffers for intermediate computations and dequantized weights.
- act_gelu -- If true, uses GELU activation; otherwise SiLU (SwiGLU).
- has_residual -- Whether to add the residual connection.
- post_layernorm and post_layernorm_bias -- Optional post-MLP normalization.
- residual_fp32 -- FP32 residual precision flag.
- use_graphs -- CUDA graph capture support.
q_mlp_forward_(q_mlp, x, loras, loras_temp) executes the fused MLP forward pass in-place on tensor x: applies norm, gate/up projection, activation (SiLU or GELU), down projection, and residual addition.
q_mlp_set_loras attaches LoRA adapter weight pairs (A and B matrices) for the gate, up, and down projections.
make_q_moe_mlp constructs a Mixture-of-Experts MLP module with additional parameters:
- gate -- The router/gating tensor for expert selection.
- num_experts and num_experts_per_token -- MoE configuration.
- w1, w2, w3 -- Vectors of quantized weight handles, one per expert.
- temp_gathered_state and temp_logits -- Additional scratch buffers for expert routing.
q_moe_mlp_forward_(q_moe_mlp, x) executes the MoE MLP forward pass: computes router logits, selects top-K experts per token, gathers states, runs per-expert MLPs, and combines outputs.
tp_mlp_forward_ is the tensor-parallel MLP variant that distributes the gate/up/down projections across multiple devices, with broadcasting and gathering managed by the ExtTPContext.
Usage
This API is used by the Python ExLlamaV2MLP and ExLlamaV2MoEMLP layer classes. Modules are constructed once during model loading, and forward functions are called on each inference step. The fused design avoids multiple kernel launches and intermediate memory allocations.
Code Reference
Source Location
- Repository: Turboderp_org_Exllamav2
- File: exllamav2/exllamav2_ext/ext_qmlp.h
- Lines: 1-110
Signature
uintptr_t make_q_mlp(
torch::Tensor layernorm,
torch::Tensor layernorm_bias,
bool layernorm_is_rms,
float norm_epsilon,
uintptr_t q_gate,
uintptr_t q_up,
uintptr_t q_down,
torch::Tensor temp_state,
torch::Tensor temp_a,
torch::Tensor temp_b,
torch::Tensor temp_dq,
int max_rows,
bool act_gelu,
bool has_residual,
torch::Tensor post_layernorm,
torch::Tensor post_layernorm_bias,
bool residual_fp32,
bool use_graphs
);
void q_mlp_forward_(
uintptr_t q_mlp,
torch::Tensor x,
const std::vector<uintptr_t>& loras,
torch::Tensor loras_temp
);
int q_mlp_set_loras(
uintptr_t q_mlp,
std::unordered_map<uintptr_t, torch::Tensor>& gate_proj_lora_a,
std::unordered_map<uintptr_t, torch::Tensor>& gate_proj_lora_b,
std::unordered_map<uintptr_t, torch::Tensor>& up_proj_lora_a,
std::unordered_map<uintptr_t, torch::Tensor>& up_proj_lora_b,
std::unordered_map<uintptr_t, torch::Tensor>& down_proj_lora_a,
std::unordered_map<uintptr_t, torch::Tensor>& down_proj_lora_b
);
uintptr_t make_q_moe_mlp(
torch::Tensor layernorm,
torch::Tensor layernorm_bias,
bool layernorm_is_rms,
float norm_epsilon,
torch::Tensor gate,
int num_experts,
int num_experts_per_token,
const std::vector<uintptr_t>& w1,
const std::vector<uintptr_t>& w2,
const std::vector<uintptr_t>& w3,
torch::Tensor temp_state,
torch::Tensor temp_gathered_state,
torch::Tensor temp_a,
torch::Tensor temp_b,
torch::Tensor temp_logits,
torch::Tensor temp_dq,
int max_rows,
bool act_gelu
);
void q_moe_mlp_forward_(
uintptr_t q_moe_mlp,
torch::Tensor x
);
void tp_mlp_forward_(
uintptr_t tp_context,
torch::Tensor hidden_states,
const std::vector<torch::Tensor> &temp_bc0,
const std::vector<torch::Tensor> &temp_bc1,
const std::vector<torch::Tensor> &temp_bc2,
const std::vector<torch::Tensor> &temp_gate,
const std::vector<torch::Tensor> &temp_up,
const std::vector<torch::Tensor> &temp_down,
const std::vector<torch::Tensor> &pre_layernorm,
float norm_epsilon,
const std::vector<uintptr_t> &gate,
const std::vector<uintptr_t> &up,
const std::vector<uintptr_t> &down,
bool act_gelu
);
Import
from exllamav2 import exllamav2_ext as ext_c
# Construct quantized MLP module
handle = ext_c.make_q_mlp(
layernorm, layernorm_bias, layernorm_is_rms, norm_epsilon,
q_gate, q_up, q_down,
temp_state, temp_a, temp_b, temp_dq,
max_rows, act_gelu, has_residual,
post_layernorm, post_layernorm_bias,
residual_fp32, use_graphs
)
I/O Contract
| Function | Key Parameters | Output | Description |
|---|---|---|---|
| make_q_mlp | Norm weights, quantized gate/up/down, scratch buffers, config flags | uintptr_t handle | Constructs fused MLP module; returns opaque handle |
| q_mlp_forward_ | q_mlp handle, input x, LoRA handles | x modified in-place | In-place MLP: norm -> gate/up -> activation -> down -> residual |
| q_mlp_set_loras | q_mlp handle, 6 LoRA weight maps (A+B for gate/up/down) | int (count) | Attaches LoRA adapters to all three projections |
| make_q_moe_mlp | Norm weights, router gate, per-expert w1/w2/w3, config | uintptr_t handle | Constructs MoE MLP module; returns opaque handle |
| q_moe_mlp_forward_ | q_moe_mlp handle, input x | x modified in-place | In-place MoE forward: route -> expert MLPs -> combine |
| tp_mlp_forward_ | tp_context, per-device tensors, norm/gate/up/down handles | hidden_states modified in-place | Tensor-parallel MLP across multiple devices |
Usage Examples
from exllamav2 import exllamav2_ext as ext_c
# Standard MLP forward pass (in-place)
ext_c.q_mlp_forward_(mlp_handle, x, loras, loras_temp)
# MoE MLP forward pass (in-place)
ext_c.q_moe_mlp_forward_(moe_handle, x)
# Tensor-parallel MLP forward
ext_c.tp_mlp_forward_(
tp_ctx, hidden_states,
temp_bc0, temp_bc1, temp_bc2,
temp_gate, temp_up, temp_down,
pre_layernorm, norm_epsilon,
gate_handles, up_handles, down_handles,
act_gelu
)
Related Pages
- Implementation:Turboderp_org_Exllamav2_Ext_QAttn_H -- Quantized attention module with similar construction pattern
- Implementation:Turboderp_org_Exllamav2_Ext_Norm -- Normalization operations used within the MLP
- Implementation:Turboderp_org_Exllamav2_Ext_TP_H -- Tensor parallelism context for tp_mlp_forward_