Jump to content

Connect SuperML | Leeroopedia MCP: Equip your AI agents with best practices, code verification, and debugging knowledge. Powered by Leeroo — building Organizational Superintelligence. Contact us at founders@leeroo.com.

Implementation:InternLM Lmdeploy GptKernels

From Leeroopedia


Knowledge Sources
Domains GPU_Kernels, Transformer
Last Updated 2026-02-07 15:00 GMT

Overview

Collection of CUDA kernel launch functions for GPT-style transformer preprocessing operations including embedding lookup, input tiling, context deduplication, padding management, and tensor transposition.

Description

This header declares the kernel invocation functions for GPT transformer input processing. Key operations include: invokeInputIdsEmbeddingLookupPosEncoding() for combined token embedding and positional encoding lookup with optional prompt tuning; invokeBuildDecoderAttentionMask() for constructing causal attention masks; invokeLookupHiddenStateOfLastToken() for extracting the final hidden state; invokeTileGptInputs() and invokeTileGptPromptInputs() for beam-search input replication; invokeFindContextDups() and invokeCompactInputs() for deduplicating shared context prefixes; invokeUpdatePaddingCount() and invokeMaskPaddingTokens() for padding handling; and invokeTransposeAxis01(), invokeTranspose2D() for tensor reshaping. The pPromptTuningParam struct encapsulates prompt-tuning configuration.

Usage

Use these kernels during the preprocessing and postprocessing stages of GPT-style inference, including context phase input preparation, beam search replication, and output extraction.

Code Reference

Source Location

Signature

template<typename T>
void invokeInputIdsEmbeddingLookupPosEncoding(
    T* from_tensor, int* output_ids, const T* embedding_table, const T* pos_table,
    pPromptTuningParam<T> prompt_param, const int* input_ids,
    const int start_step, const int length, const int max_length,
    const int batch_size, const int hidden_units, cudaStream_t stream);

template<typename T>
void invokeBuildDecoderAttentionMask(
    T* attention_mask, const int* sequence_lengths, const int* prefix_prompt_lengths,
    const int batch_size, const int max_seq_len, const int max_prompt_length,
    cudaStream_t stream);

void invokeTileGptInputs(int* tiled_input_ids, int* tiled_input_lengths,
    const int* input_ids, const int* input_lengths,
    const int batch_size, const int beam_width, const int max_input_length,
    cudaStream_t stream);

void invokeEmbeddingLookup(Ref<Tensor> out_, const Buffer_<int>& token_ids,
    const Tensor& embedding_table, cudaStream_t st);

Import

#include "src/turbomind/kernels/gpt_kernels.h"

I/O Contract

Inputs

Name Type Required Description
input_ids const int* Yes Token IDs for the input sequence
embedding_table const T* Yes Token embedding weight matrix
pos_table const T* No Positional encoding table (may be nullptr)
batch_size int Yes Number of sequences in the batch
hidden_units int Yes Hidden dimension size
stream cudaStream_t Yes CUDA stream for async execution

Outputs

Name Type Description
from_tensor T* Embedded + position-encoded output tensor
output_ids int* Processed token IDs (may be remapped)
attention_mask T* Causal attention mask

Usage Examples

using namespace turbomind;

// Embedding lookup with positional encoding
invokeInputIdsEmbeddingLookupPosEncoding(
    embedded_output, output_ids, embed_table, pos_table,
    prompt_param, input_ids, 0, seq_len, max_seq_len,
    batch_size, hidden_dim, stream);

// Tile inputs for beam search
invokeTileGptInputs(tiled_ids, tiled_lengths,
    input_ids, input_lengths, batch_size, beam_width, max_len, stream);

Related Pages

Page Connections

Double-click a node to navigate. Hold to expand connections.
Principle
Implementation
Heuristic
Environment