Jump to content

Connect SuperML | Leeroopedia MCP: Equip your AI agents with best practices, code verification, and debugging knowledge. Powered by Leeroo — building Organizational Superintelligence. Contact us at founders@leeroo.com.

Implementation:Turboderp org Exllamav2 Ext QMLP H

From Leeroopedia
Knowledge Sources
Domains MLP, Quantization, C_Extension
Last Updated 2026-02-15 00:00 GMT

Overview

C++ header declaring the API for ExLlamaV2's fused quantized MLP and Mixture-of-Experts (MoE) MLP modules, including construction, forward pass, LoRA support, and tensor-parallel variants.

Description

ext_qmlp.h defines the interface for two quantized feed-forward network subsystems:

make_q_mlp constructs a standard quantized MLP module. It accepts parameters including:

  • layernorm and layernorm_bias -- Pre-MLP normalization weights.
  • layernorm_is_rms -- Flag selecting RMS norm vs standard layer norm.
  • q_gate, q_up, q_down -- Opaque handles to quantized weight matrices for the gate, up-projection, and down-projection layers.
  • temp_state, temp_a, temp_b, temp_dq -- Scratch buffers for intermediate computations and dequantized weights.
  • act_gelu -- If true, uses GELU activation; otherwise SiLU (SwiGLU).
  • has_residual -- Whether to add the residual connection.
  • post_layernorm and post_layernorm_bias -- Optional post-MLP normalization.
  • residual_fp32 -- FP32 residual precision flag.
  • use_graphs -- CUDA graph capture support.

q_mlp_forward_(q_mlp, x, loras, loras_temp) executes the fused MLP forward pass in-place on tensor x: applies norm, gate/up projection, activation (SiLU or GELU), down projection, and residual addition.

q_mlp_set_loras attaches LoRA adapter weight pairs (A and B matrices) for the gate, up, and down projections.

make_q_moe_mlp constructs a Mixture-of-Experts MLP module with additional parameters:

  • gate -- The router/gating tensor for expert selection.
  • num_experts and num_experts_per_token -- MoE configuration.
  • w1, w2, w3 -- Vectors of quantized weight handles, one per expert.
  • temp_gathered_state and temp_logits -- Additional scratch buffers for expert routing.

q_moe_mlp_forward_(q_moe_mlp, x) executes the MoE MLP forward pass: computes router logits, selects top-K experts per token, gathers states, runs per-expert MLPs, and combines outputs.

tp_mlp_forward_ is the tensor-parallel MLP variant that distributes the gate/up/down projections across multiple devices, with broadcasting and gathering managed by the ExtTPContext.

Usage

This API is used by the Python ExLlamaV2MLP and ExLlamaV2MoEMLP layer classes. Modules are constructed once during model loading, and forward functions are called on each inference step. The fused design avoids multiple kernel launches and intermediate memory allocations.

Code Reference

Source Location

Signature

uintptr_t make_q_mlp(
    torch::Tensor layernorm,
    torch::Tensor layernorm_bias,
    bool layernorm_is_rms,
    float norm_epsilon,
    uintptr_t q_gate,
    uintptr_t q_up,
    uintptr_t q_down,
    torch::Tensor temp_state,
    torch::Tensor temp_a,
    torch::Tensor temp_b,
    torch::Tensor temp_dq,
    int max_rows,
    bool act_gelu,
    bool has_residual,
    torch::Tensor post_layernorm,
    torch::Tensor post_layernorm_bias,
    bool residual_fp32,
    bool use_graphs
);

void q_mlp_forward_(
    uintptr_t q_mlp,
    torch::Tensor x,
    const std::vector<uintptr_t>& loras,
    torch::Tensor loras_temp
);

int q_mlp_set_loras(
    uintptr_t q_mlp,
    std::unordered_map<uintptr_t, torch::Tensor>& gate_proj_lora_a,
    std::unordered_map<uintptr_t, torch::Tensor>& gate_proj_lora_b,
    std::unordered_map<uintptr_t, torch::Tensor>& up_proj_lora_a,
    std::unordered_map<uintptr_t, torch::Tensor>& up_proj_lora_b,
    std::unordered_map<uintptr_t, torch::Tensor>& down_proj_lora_a,
    std::unordered_map<uintptr_t, torch::Tensor>& down_proj_lora_b
);

uintptr_t make_q_moe_mlp(
    torch::Tensor layernorm,
    torch::Tensor layernorm_bias,
    bool layernorm_is_rms,
    float norm_epsilon,
    torch::Tensor gate,
    int num_experts,
    int num_experts_per_token,
    const std::vector<uintptr_t>& w1,
    const std::vector<uintptr_t>& w2,
    const std::vector<uintptr_t>& w3,
    torch::Tensor temp_state,
    torch::Tensor temp_gathered_state,
    torch::Tensor temp_a,
    torch::Tensor temp_b,
    torch::Tensor temp_logits,
    torch::Tensor temp_dq,
    int max_rows,
    bool act_gelu
);

void q_moe_mlp_forward_(
    uintptr_t q_moe_mlp,
    torch::Tensor x
);

void tp_mlp_forward_(
    uintptr_t tp_context,
    torch::Tensor hidden_states,
    const std::vector<torch::Tensor> &temp_bc0,
    const std::vector<torch::Tensor> &temp_bc1,
    const std::vector<torch::Tensor> &temp_bc2,
    const std::vector<torch::Tensor> &temp_gate,
    const std::vector<torch::Tensor> &temp_up,
    const std::vector<torch::Tensor> &temp_down,
    const std::vector<torch::Tensor> &pre_layernorm,
    float norm_epsilon,
    const std::vector<uintptr_t> &gate,
    const std::vector<uintptr_t> &up,
    const std::vector<uintptr_t> &down,
    bool act_gelu
);

Import

from exllamav2 import exllamav2_ext as ext_c

# Construct quantized MLP module
handle = ext_c.make_q_mlp(
    layernorm, layernorm_bias, layernorm_is_rms, norm_epsilon,
    q_gate, q_up, q_down,
    temp_state, temp_a, temp_b, temp_dq,
    max_rows, act_gelu, has_residual,
    post_layernorm, post_layernorm_bias,
    residual_fp32, use_graphs
)

I/O Contract

Function Key Parameters Output Description
make_q_mlp Norm weights, quantized gate/up/down, scratch buffers, config flags uintptr_t handle Constructs fused MLP module; returns opaque handle
q_mlp_forward_ q_mlp handle, input x, LoRA handles x modified in-place In-place MLP: norm -> gate/up -> activation -> down -> residual
q_mlp_set_loras q_mlp handle, 6 LoRA weight maps (A+B for gate/up/down) int (count) Attaches LoRA adapters to all three projections
make_q_moe_mlp Norm weights, router gate, per-expert w1/w2/w3, config uintptr_t handle Constructs MoE MLP module; returns opaque handle
q_moe_mlp_forward_ q_moe_mlp handle, input x x modified in-place In-place MoE forward: route -> expert MLPs -> combine
tp_mlp_forward_ tp_context, per-device tensors, norm/gate/up/down handles hidden_states modified in-place Tensor-parallel MLP across multiple devices

Usage Examples

from exllamav2 import exllamav2_ext as ext_c

# Standard MLP forward pass (in-place)
ext_c.q_mlp_forward_(mlp_handle, x, loras, loras_temp)

# MoE MLP forward pass (in-place)
ext_c.q_moe_mlp_forward_(moe_handle, x)

# Tensor-parallel MLP forward
ext_c.tp_mlp_forward_(
    tp_ctx, hidden_states,
    temp_bc0, temp_bc1, temp_bc2,
    temp_gate, temp_up, temp_down,
    pre_layernorm, norm_epsilon,
    gate_handles, up_handles, down_handles,
    act_gelu
)

Related Pages

Page Connections

Double-click a node to navigate. Hold to expand connections.
Principle
Implementation
Heuristic
Environment