Implementation:Microsoft DeepSpeedExamples Domino Language Model
| Knowledge Sources | |
|---|---|
| Domains | Natural Language Processing, Distributed Training |
| Last Updated | 2026-02-07 12:00 GMT |
Overview
Domino-adapted transformer language model built on Megatron-LM that integrates DeepSpeed's DominoTransformer for intra-layer communication overlap in distributed training.
Description
This module implements a transformer-based language model adapted from Megatron-LM's language_model.py for the DeepSpeed Domino distributed training system. The key innovation is the integration of DominoTransformer from deepspeed.runtime.domino.transformer as the encoder backbone, which enables overlapping intra-layer communication with computation during distributed training to reduce communication overhead.
The module provides three main classes: Embedding for language model embeddings (word, position, and token-type embeddings with parallel vocabulary support), Pooler for extracting sequence-level representations via a linear+tanh transformation, and TransformerLanguageModel as the main model class that assembles the full language model. The TransformerLanguageModel supports rotary position embeddings (RoPE), learned absolute positional embeddings, optional encoder-decoder architecture, and configurable pre/post-processing stages for pipeline parallelism.
The module also provides parallel_lm_logits for computing language model logits with tensor-model-parallel support and get_language_model as a factory function for building the language model with proper initialization. All components integrate with Megatron's model parallel utilities (mpu) for tensor-parallel and pipeline-parallel distributed training.
Usage
Use this module as part of the DeepSpeed-Domino training example to build transformer language models that benefit from Domino's intra-layer overlap optimization. It is designed to work within the Megatron-LM framework with DeepSpeed integration for large-scale distributed pretraining of GPT-style language models.
Code Reference
Source Location
- Repository: Microsoft_DeepSpeedExamples
- File: training/DeepSpeed-Domino/domino/language_model.py
- Lines: 1-602
Signature
def parallel_lm_logits(input_, word_embeddings_weight, parallel_output, bias=None):
def get_language_model(config, num_tokentypes, add_pooler,
encoder_attn_mask_type, add_encoder=True,
add_decoder=False, decoder_attn_mask_type=AttnMaskType.causal,
pre_process=True, post_process=True):
class Pooler(MegatronModule):
def __init__(self, hidden_size, init_method):
class Embedding(MegatronModule):
def __init__(self, hidden_size, vocab_size, max_sequence_length,
embedding_dropout_prob, config, num_tokentypes=0,
embedding_weights_in_fp32=False):
class TransformerLanguageModel(MegatronModule):
def __init__(self, config, encoder_attn_mask_type, num_tokentypes=0,
add_encoder=True, add_decoder=False,
decoder_attn_mask_type=AttnMaskType.causal,
add_pooler=False, pre_process=True, post_process=True):
Import
from domino.language_model import get_language_model, TransformerLanguageModel, parallel_lm_logits
I/O Contract
Inputs
| Name | Type | Required | Description |
|---|---|---|---|
| config | TransformerConfig | Yes | Transformer configuration with init_method, hidden_size, num_layers, etc. |
| num_tokentypes | int | Yes | Number of token types (0 to disable token-type embeddings) |
| add_pooler | bool | Yes | Whether to add a pooler layer for sequence-level tasks |
| encoder_attn_mask_type | AttnMaskType | Yes | Attention mask type (e.g., padding, causal) |
| pre_process | bool | No | Whether this stage handles embedding (pipeline parallelism) |
| post_process | bool | No | Whether this stage handles output projection (pipeline parallelism) |
| input_ids | torch.LongTensor | Yes | Token indices for the input sequence |
| position_ids | torch.LongTensor | Yes | Position indices for positional embeddings |
Outputs
| Name | Type | Description |
|---|---|---|
| language_model | TransformerLanguageModel | The constructed language model instance |
| language_model_key | str | Checkpoint key string ('language_model') |
| lm_output | torch.Tensor | Encoder hidden states of shape [seq_len, batch, hidden_size] |
| pooled_output | torch.Tensor | Pooled representation of shape [batch, hidden_size] (if pooler enabled) |
Usage Examples
from domino.language_model import get_language_model
from megatron.model.enums import AttnMaskType
# Build language model with Domino transformer
language_model, lm_key = get_language_model(
config=config,
num_tokentypes=0,
add_pooler=False,
encoder_attn_mask_type=AttnMaskType.causal,
add_encoder=True,
add_decoder=False,
pre_process=True,
post_process=True
)
# Forward pass
lm_output = language_model(
enc_input_ids=input_ids,
enc_position_ids=position_ids,
enc_attn_mask=attention_mask
)