Jump to content

Connect SuperML | Leeroopedia MCP: Equip your AI agents with best practices, code verification, and debugging knowledge. Powered by Leeroo — building Organizational Superintelligence. Contact us at founders@leeroo.com.

Implementation:NVIDIA NeMo Curator TokenizerStage

From Leeroopedia
Revision as of 10:48, 27 September 2026 by Agent (talk | contribs) (Sync from local file)
(diff) ← Older revision | Latest revision (diff) | Newer revision → (diff)
Knowledge Sources
Domains NLP, Tokenization, HuggingFace, Data Pipeline
Last Updated 2026-02-14 00:00 GMT

Overview

Processing stage that tokenizes text fields in a DocumentBatch using a HuggingFace AutoTokenizer, preparing inputs for downstream GPU model inference stages.

Description

TokenizerStage extends ProcessingStage[DocumentBatch, DocumentBatch] and serves as the essential preprocessing bridge between raw text data and GPU model inference. It loads a HuggingFace AutoTokenizer on setup and applies batch tokenization with configurable padding, truncation, and sequence length limits.

Key behaviors:

  • Model download: On node setup, downloads the model files via snapshot_download. On worker setup, loads the tokenizer from the local cache.
  • Text truncation: Optionally truncates text to max_chars characters before tokenization.
  • Tokenization: Uses batch_encode_plus with padding to max_length, truncation enabled, and special tokens added. Returns input_ids and attention_mask as numpy arrays.
  • Sequence length handling: If max_seq_length is not specified, uses the tokenizer's model_max_length. Includes a guard against a known HuggingFace bug where some models report extremely large max lengths (> 100,000), falling back to max_position_embeddings from the model config.
  • Length-based sorting: Optionally sorts rows by token length to improve GPU batching efficiency. Adds a _curator_seq_order column to preserve and later restore original ordering.

Usage

Use TokenizerStage as a preprocessing step before any ModelStage-based inference. It is automatically included when using EmbeddingCreatorStage, but can also be used independently for custom pipelines that need tokenized data.

Code Reference

Source Location

  • Repository: NeMo-Curator
  • File: nemo_curator/stages/text/models/tokenizer.py
  • Lines: 1-170

Signature

class TokenizerStage(ProcessingStage[DocumentBatch, DocumentBatch]):
    def __init__(
        self,
        model_identifier: str,
        cache_dir: str | None = None,
        hf_token: str | None = None,
        text_field: str = "text",
        max_chars: int | None = None,
        max_seq_length: int | None = None,
        padding_side: Literal["left", "right"] = "right",
        sort_by_length: bool = True,
        unk_token: bool = False,
    ): ...

Import

from nemo_curator.stages.text.models.tokenizer import TokenizerStage

I/O Contract

Inputs

Name Type Required Description
model_identifier str Yes HuggingFace model identifier or local path for the tokenizer
cache_dir str or None No Directory for caching downloaded model/tokenizer files (default: None)
hf_token str or None No HuggingFace authentication token for gated models (default: None)
text_field str No Name of the text column in the input DocumentBatch (default: "text")
max_chars int or None No Maximum character count for text truncation before tokenization (default: None)
max_seq_length int or None No Maximum token sequence length; if None, derived from the model config (default: None)
padding_side Literal["left", "right"] No Side on which to pad tokenized sequences (default: "right")
sort_by_length bool No Whether to sort rows by token length for GPU batching efficiency (default: True)
unk_token bool No If True, sets the pad token to the tokenizer's unknown token (default: False)

Stage I/O Specification

Method Returns
inputs() (["data"], [text_field])
outputs() (["data"], [text_field, "_curator_input_ids", "_curator_attention_mask"] + optionally ["_curator_seq_order"])

Outputs

Name Type Description
DocumentBatch DocumentBatch Input data augmented with _curator_input_ids (list of int per row), _curator_attention_mask (list of int per row), and optionally _curator_seq_order (int) columns

Usage Examples

Basic Usage

from nemo_curator.stages.text.models.tokenizer import TokenizerStage

tokenizer = TokenizerStage(
    model_identifier="sentence-transformers/all-MiniLM-L6-v2",
    text_field="text",
    max_seq_length=512,
    sort_by_length=True,
)

With Character Truncation and Custom Padding

from nemo_curator.stages.text.models.tokenizer import TokenizerStage

tokenizer = TokenizerStage(
    model_identifier="intfloat/e5-large-v2",
    text_field="content",
    max_chars=5000,
    padding_side="left",
    sort_by_length=False,
)

Using unk_token for Models Without a Pad Token

from nemo_curator.stages.text.models.tokenizer import TokenizerStage

# Some models (e.g., LLaMA) don't define a pad token by default
tokenizer = TokenizerStage(
    model_identifier="meta-llama/Llama-2-7b-hf",
    hf_token="hf_...",
    unk_token=True,
)

Implementation Details

Max Sequence Length Guard

Some HuggingFace models report their model_max_length as an extremely large integer (e.g. int(1e30)). The _setup method detects this case (values greater than 100,000) and falls back to AutoConfig.from_pretrained(...).max_position_embeddings to get the actual maximum position embedding size. The load_cfg method is decorated with @lru_cache(maxsize=1) to avoid redundant config loading.

Sort-by-Length Optimization

When sort_by_length is True, the stage:

  1. Computes the actual token count per row by summing the attention_mask values
  2. Adds a _curator_seq_order column with original row indices (0, 1, 2, ...)
  3. Sorts by token length using stable sort (preserving relative order of equal-length sequences)
  4. Drops the temporary _curator_token_length column

This sorting ensures that sequences of similar length are batched together during GPU inference, reducing wasted computation on padding tokens. The downstream ModelStage uses the _curator_seq_order column to restore original ordering.

Ray Actor Stage

The ray_stage_spec method returns {"is_actor_stage": True}, indicating that this stage should be scheduled as a Ray actor. This is useful because the tokenizer instance maintains state (the loaded tokenizer model) that should persist across multiple batch invocations.

Tokenization Details

The batch_encode_plus call uses:

  • max_length=self.max_seq_length - enforces the maximum sequence length
  • padding="max_length" - pads all sequences to the same length
  • return_tensors="np" - returns NumPy arrays for efficient conversion
  • truncation=True - truncates sequences exceeding max_length
  • add_special_tokens=True - adds model-specific special tokens (CLS, SEP, etc.)
  • return_token_type_ids=False - omits token type IDs (not needed for most embedding models)

Related Pages

Page Connections

Double-click a node to navigate. Hold to expand connections.
Principle
Implementation
Heuristic
Environment