Jump to content

Connect SuperML | Leeroopedia MCP: Equip your AI agents with best practices, code verification, and debugging knowledge. Powered by Leeroo — building Organizational Superintelligence. Contact us at founders@leeroo.com.

Implementation:Hiyouga LLaMA Factory KTransformers DeepSeek V3 4GPU Rules

From Leeroopedia
Revision as of 10:41, 27 September 2026 by Agent (talk | contribs) (Sync from local file)
(diff) ← Older revision | Latest revision (diff) | Newer revision → (diff)


Knowledge Sources
Domains Model Optimization, Multi-GPU Inference
Last Updated 2026-02-06 19:00 GMT

Overview

Concrete YAML configuration file for distributing DeepSeek-V3 model layers across 4 GPUs with AMX-accelerated CPU offloading provided by LLaMA Factory.

Description

This YAML file defines KTransformers optimization rules for running DeepSeek-V3-Chat in a supervised fine-tuning (SFT) configuration across 4 GPUs using Intel AMX (Advanced Matrix Extensions) for CPU-based MoE expert computation. The file contains a list of match-and-replace rules that map specific model components (identified by regex patterns on module names) to their KTransformers operator replacements, with explicit GPU device assignments.

The 61 transformer layers (0-60) of DeepSeek-V3 are partitioned across four GPUs:

  • GPU 0 (cuda:0): Layers 0-14
  • GPU 1 (cuda:1): Layers 15-29
  • GPU 2 (cuda:2): Layers 30-44
  • GPU 3 (cuda:3): Layers 45-60

Each rule section handles a different component type: rotary embeddings, linear layers (excluding kv_b_proj), MLP/MoE modules, MoE gates, MoE experts, self-attention blocks, and catch-all defaults. The MoE experts use CPU offloading with AMXInt8 backend for generation and GPU-based KExpertsTorch for prefill.

Usage

This configuration is used when running DeepSeek-V3-Chat inference or SFT with KTransformers on a 4-GPU system. It is referenced via the kt_optimize_rules parameter in a KTransformers training or inference YAML configuration file. The file is loaded by the KTransformers runtime to determine how to replace standard PyTorch modules with optimized KTransformers operators.

Code Reference

Source Location

Signature

# Top-level structure: list of match/replace rule dictionaries
- match:
    name: "<regex_pattern>"
    class: <original_class>  # optional
  replace:
    class: <replacement_class_or_default>
    kwargs:
      generate_device: "<device>"
      prefill_device: "<device>"
      # additional operator-specific kwargs
  recursive: False  # optional

Import

# Referenced in a KTransformers training config:
kt_optimize_rules: examples/ktransformers/kt_optimize_rules/DeepSeek-V3-Chat-sft-amx-multi-gpu-4.yaml

I/O Contract

Inputs

Name Type Required Description
match.name string (regex) Yes Regex pattern matching the full module name in the model
match.class string No Fully qualified class name of the original module to match
replace.class string Yes Replacement KTransformers operator class or "default" to keep original class
replace.kwargs dict Yes Device placement and operator configuration arguments
recursive bool No Whether to apply replacement recursively (defaults to True)

Outputs

Name Type Description
Optimized model KTransformers model DeepSeek-V3 model with layers distributed across 4 GPUs and MoE experts offloaded to CPU with AMX acceleration

Component Replacement Rules

Component Original Class Replacement Class Notes
Embedding model.embed_tokens default CPU for both generate and prefill
Rotary Embedding DeepseekV3RotaryEmbedding YarnRotaryEmbeddingV3 Per-GPU assignment by layer range
Linear (excl. kv_b_proj) torch.nn.Linear KTransformersLinear KLinearTorch for both phases
MLP (MoE) DeepseekV3MoE KDeepseekV3MoE Per-GPU assignment
MoE Gate MoEGate KMoEGate Per-GPU assignment
MoE Experts (experts) KTransformersExperts GPU prefill, CPU generate with AMXInt8
Self-Attention (self_attn) KDeepseekV2Attention absorb_for_prefill=False
Overall Model model KDeepseekV2Model Defines transfer_map for layer boundaries
LM Head torch.nn.Linear KTransformersLinear On GPU 3
Model Norm (model.norm) default On GPU 3

Usage Examples

# Example KTransformers training configuration referencing this rules file
model_name_or_path: deepseek-ai/DeepSeek-V3-Chat
stage: sft
do_train: true
finetuning_type: lora
use_kt: true
kt_optimize_rules: examples/ktransformers/kt_optimize_rules/DeepSeek-V3-Chat-sft-amx-multi-gpu-4.yaml

Related Pages

Page Connections

Double-click a node to navigate. Hold to expand connections.
Principle
Implementation
Heuristic
Environment