Jump to content

Connect SuperML | Leeroopedia MCP: Equip your AI agents with best practices, code verification, and debugging knowledge. Powered by Leeroo — building Organizational Superintelligence. Contact us at founders@leeroo.com.

Implementation:Microsoft DeepSpeedExamples Domino Language Model

From Leeroopedia


Knowledge Sources
Domains Natural Language Processing, Distributed Training
Last Updated 2026-02-07 12:00 GMT

Overview

Domino-adapted transformer language model built on Megatron-LM that integrates DeepSpeed's DominoTransformer for intra-layer communication overlap in distributed training.

Description

This module implements a transformer-based language model adapted from Megatron-LM's language_model.py for the DeepSpeed Domino distributed training system. The key innovation is the integration of DominoTransformer from deepspeed.runtime.domino.transformer as the encoder backbone, which enables overlapping intra-layer communication with computation during distributed training to reduce communication overhead.

The module provides three main classes: Embedding for language model embeddings (word, position, and token-type embeddings with parallel vocabulary support), Pooler for extracting sequence-level representations via a linear+tanh transformation, and TransformerLanguageModel as the main model class that assembles the full language model. The TransformerLanguageModel supports rotary position embeddings (RoPE), learned absolute positional embeddings, optional encoder-decoder architecture, and configurable pre/post-processing stages for pipeline parallelism.

The module also provides parallel_lm_logits for computing language model logits with tensor-model-parallel support and get_language_model as a factory function for building the language model with proper initialization. All components integrate with Megatron's model parallel utilities (mpu) for tensor-parallel and pipeline-parallel distributed training.

Usage

Use this module as part of the DeepSpeed-Domino training example to build transformer language models that benefit from Domino's intra-layer overlap optimization. It is designed to work within the Megatron-LM framework with DeepSpeed integration for large-scale distributed pretraining of GPT-style language models.

Code Reference

Source Location

Signature

def parallel_lm_logits(input_, word_embeddings_weight, parallel_output, bias=None):

def get_language_model(config, num_tokentypes, add_pooler,
                       encoder_attn_mask_type, add_encoder=True,
                       add_decoder=False, decoder_attn_mask_type=AttnMaskType.causal,
                       pre_process=True, post_process=True):

class Pooler(MegatronModule):
    def __init__(self, hidden_size, init_method):

class Embedding(MegatronModule):
    def __init__(self, hidden_size, vocab_size, max_sequence_length,
                 embedding_dropout_prob, config, num_tokentypes=0,
                 embedding_weights_in_fp32=False):

class TransformerLanguageModel(MegatronModule):
    def __init__(self, config, encoder_attn_mask_type, num_tokentypes=0,
                 add_encoder=True, add_decoder=False,
                 decoder_attn_mask_type=AttnMaskType.causal,
                 add_pooler=False, pre_process=True, post_process=True):

Import

from domino.language_model import get_language_model, TransformerLanguageModel, parallel_lm_logits

I/O Contract

Inputs

Name Type Required Description
config TransformerConfig Yes Transformer configuration with init_method, hidden_size, num_layers, etc.
num_tokentypes int Yes Number of token types (0 to disable token-type embeddings)
add_pooler bool Yes Whether to add a pooler layer for sequence-level tasks
encoder_attn_mask_type AttnMaskType Yes Attention mask type (e.g., padding, causal)
pre_process bool No Whether this stage handles embedding (pipeline parallelism)
post_process bool No Whether this stage handles output projection (pipeline parallelism)
input_ids torch.LongTensor Yes Token indices for the input sequence
position_ids torch.LongTensor Yes Position indices for positional embeddings

Outputs

Name Type Description
language_model TransformerLanguageModel The constructed language model instance
language_model_key str Checkpoint key string ('language_model')
lm_output torch.Tensor Encoder hidden states of shape [seq_len, batch, hidden_size]
pooled_output torch.Tensor Pooled representation of shape [batch, hidden_size] (if pooler enabled)

Usage Examples

from domino.language_model import get_language_model
from megatron.model.enums import AttnMaskType

# Build language model with Domino transformer
language_model, lm_key = get_language_model(
    config=config,
    num_tokentypes=0,
    add_pooler=False,
    encoder_attn_mask_type=AttnMaskType.causal,
    add_encoder=True,
    add_decoder=False,
    pre_process=True,
    post_process=True
)

# Forward pass
lm_output = language_model(
    enc_input_ids=input_ids,
    enc_position_ids=position_ids,
    enc_attn_mask=attention_mask
)

Related Pages

Page Connections

Double-click a node to navigate. Hold to expand connections.
Principle
Implementation
Heuristic
Environment