Implementation:Hiyouga LLaMA Factory KTransformers DeepSeek V3 4GPU Rules
| Knowledge Sources | |
|---|---|
| Domains | Model Optimization, Multi-GPU Inference |
| Last Updated | 2026-02-06 19:00 GMT |
Overview
Concrete YAML configuration file for distributing DeepSeek-V3 model layers across 4 GPUs with AMX-accelerated CPU offloading provided by LLaMA Factory.
Description
This YAML file defines KTransformers optimization rules for running DeepSeek-V3-Chat in a supervised fine-tuning (SFT) configuration across 4 GPUs using Intel AMX (Advanced Matrix Extensions) for CPU-based MoE expert computation. The file contains a list of match-and-replace rules that map specific model components (identified by regex patterns on module names) to their KTransformers operator replacements, with explicit GPU device assignments.
The 61 transformer layers (0-60) of DeepSeek-V3 are partitioned across four GPUs:
- GPU 0 (cuda:0): Layers 0-14
- GPU 1 (cuda:1): Layers 15-29
- GPU 2 (cuda:2): Layers 30-44
- GPU 3 (cuda:3): Layers 45-60
Each rule section handles a different component type: rotary embeddings, linear layers (excluding kv_b_proj), MLP/MoE modules, MoE gates, MoE experts, self-attention blocks, and catch-all defaults. The MoE experts use CPU offloading with AMXInt8 backend for generation and GPU-based KExpertsTorch for prefill.
Usage
This configuration is used when running DeepSeek-V3-Chat inference or SFT with KTransformers on a 4-GPU system. It is referenced via the kt_optimize_rules parameter in a KTransformers training or inference YAML configuration file. The file is loaded by the KTransformers runtime to determine how to replace standard PyTorch modules with optimized KTransformers operators.
Code Reference
Source Location
- Repository: Hiyouga_LLaMA_Factory
- File: examples/ktransformers/kt_optimize_rules/DeepSeek-V3-Chat-sft-amx-multi-gpu-4.yaml
- Lines: 1-392
Signature
# Top-level structure: list of match/replace rule dictionaries
- match:
name: "<regex_pattern>"
class: <original_class> # optional
replace:
class: <replacement_class_or_default>
kwargs:
generate_device: "<device>"
prefill_device: "<device>"
# additional operator-specific kwargs
recursive: False # optional
Import
# Referenced in a KTransformers training config:
kt_optimize_rules: examples/ktransformers/kt_optimize_rules/DeepSeek-V3-Chat-sft-amx-multi-gpu-4.yaml
I/O Contract
Inputs
| Name | Type | Required | Description |
|---|---|---|---|
| match.name | string (regex) | Yes | Regex pattern matching the full module name in the model |
| match.class | string | No | Fully qualified class name of the original module to match |
| replace.class | string | Yes | Replacement KTransformers operator class or "default" to keep original class |
| replace.kwargs | dict | Yes | Device placement and operator configuration arguments |
| recursive | bool | No | Whether to apply replacement recursively (defaults to True) |
Outputs
| Name | Type | Description |
|---|---|---|
| Optimized model | KTransformers model | DeepSeek-V3 model with layers distributed across 4 GPUs and MoE experts offloaded to CPU with AMX acceleration |
Component Replacement Rules
| Component | Original Class | Replacement Class | Notes |
|---|---|---|---|
| Embedding | model.embed_tokens | default | CPU for both generate and prefill |
| Rotary Embedding | DeepseekV3RotaryEmbedding | YarnRotaryEmbeddingV3 | Per-GPU assignment by layer range |
| Linear (excl. kv_b_proj) | torch.nn.Linear | KTransformersLinear | KLinearTorch for both phases |
| MLP (MoE) | DeepseekV3MoE | KDeepseekV3MoE | Per-GPU assignment |
| MoE Gate | MoEGate | KMoEGate | Per-GPU assignment |
| MoE Experts | (experts) | KTransformersExperts | GPU prefill, CPU generate with AMXInt8 |
| Self-Attention | (self_attn) | KDeepseekV2Attention | absorb_for_prefill=False |
| Overall Model | model | KDeepseekV2Model | Defines transfer_map for layer boundaries |
| LM Head | torch.nn.Linear | KTransformersLinear | On GPU 3 |
| Model Norm | (model.norm) | default | On GPU 3 |
Usage Examples
# Example KTransformers training configuration referencing this rules file
model_name_or_path: deepseek-ai/DeepSeek-V3-Chat
stage: sft
do_train: true
finetuning_type: lora
use_kt: true
kt_optimize_rules: examples/ktransformers/kt_optimize_rules/DeepSeek-V3-Chat-sft-amx-multi-gpu-4.yaml
Related Pages
- Implementation:Hiyouga_LLaMA_Factory_Constants - Engine name and model configuration constants
- Implementation:Hiyouga_LLaMA_Factory_Misc_Utils - use_kt() helper function