Implementation:InternLM Lmdeploy GptKernels
| Knowledge Sources | |
|---|---|
| Domains | GPU_Kernels, Transformer |
| Last Updated | 2026-02-07 15:00 GMT |
Overview
Collection of CUDA kernel launch functions for GPT-style transformer preprocessing operations including embedding lookup, input tiling, context deduplication, padding management, and tensor transposition.
Description
This header declares the kernel invocation functions for GPT transformer input processing. Key operations include: invokeInputIdsEmbeddingLookupPosEncoding() for combined token embedding and positional encoding lookup with optional prompt tuning; invokeBuildDecoderAttentionMask() for constructing causal attention masks; invokeLookupHiddenStateOfLastToken() for extracting the final hidden state; invokeTileGptInputs() and invokeTileGptPromptInputs() for beam-search input replication; invokeFindContextDups() and invokeCompactInputs() for deduplicating shared context prefixes; invokeUpdatePaddingCount() and invokeMaskPaddingTokens() for padding handling; and invokeTransposeAxis01(), invokeTranspose2D() for tensor reshaping. The pPromptTuningParam struct encapsulates prompt-tuning configuration.
Usage
Use these kernels during the preprocessing and postprocessing stages of GPT-style inference, including context phase input preparation, beam search replication, and output extraction.
Code Reference
Source Location
- Repository: InternLM_Lmdeploy
- File: src/turbomind/kernels/gpt_kernels.h
Signature
template<typename T>
void invokeInputIdsEmbeddingLookupPosEncoding(
T* from_tensor, int* output_ids, const T* embedding_table, const T* pos_table,
pPromptTuningParam<T> prompt_param, const int* input_ids,
const int start_step, const int length, const int max_length,
const int batch_size, const int hidden_units, cudaStream_t stream);
template<typename T>
void invokeBuildDecoderAttentionMask(
T* attention_mask, const int* sequence_lengths, const int* prefix_prompt_lengths,
const int batch_size, const int max_seq_len, const int max_prompt_length,
cudaStream_t stream);
void invokeTileGptInputs(int* tiled_input_ids, int* tiled_input_lengths,
const int* input_ids, const int* input_lengths,
const int batch_size, const int beam_width, const int max_input_length,
cudaStream_t stream);
void invokeEmbeddingLookup(Ref<Tensor> out_, const Buffer_<int>& token_ids,
const Tensor& embedding_table, cudaStream_t st);
Import
#include "src/turbomind/kernels/gpt_kernels.h"
I/O Contract
Inputs
| Name | Type | Required | Description |
|---|---|---|---|
| input_ids | const int* | Yes | Token IDs for the input sequence |
| embedding_table | const T* | Yes | Token embedding weight matrix |
| pos_table | const T* | No | Positional encoding table (may be nullptr) |
| batch_size | int | Yes | Number of sequences in the batch |
| hidden_units | int | Yes | Hidden dimension size |
| stream | cudaStream_t | Yes | CUDA stream for async execution |
Outputs
| Name | Type | Description |
|---|---|---|
| from_tensor | T* | Embedded + position-encoded output tensor |
| output_ids | int* | Processed token IDs (may be remapped) |
| attention_mask | T* | Causal attention mask |
Usage Examples
using namespace turbomind;
// Embedding lookup with positional encoding
invokeInputIdsEmbeddingLookupPosEncoding(
embedded_output, output_ids, embed_table, pos_table,
prompt_param, input_ids, 0, seq_len, max_seq_len,
batch_size, hidden_dim, stream);
// Tile inputs for beam search
invokeTileGptInputs(tiled_ids, tiled_lengths,
input_ids, input_lengths, batch_size, beam_width, max_len, stream);