Implementation:NVIDIA NeMo Curator TokenizerStage
| Knowledge Sources | |
|---|---|
| Domains | NLP, Tokenization, HuggingFace, Data Pipeline |
| Last Updated | 2026-02-14 00:00 GMT |
Overview
Processing stage that tokenizes text fields in a DocumentBatch using a HuggingFace AutoTokenizer, preparing inputs for downstream GPU model inference stages.
Description
TokenizerStage extends ProcessingStage[DocumentBatch, DocumentBatch] and serves as the essential preprocessing bridge between raw text data and GPU model inference. It loads a HuggingFace AutoTokenizer on setup and applies batch tokenization with configurable padding, truncation, and sequence length limits.
Key behaviors:
- Model download: On node setup, downloads the model files via snapshot_download. On worker setup, loads the tokenizer from the local cache.
- Text truncation: Optionally truncates text to max_chars characters before tokenization.
- Tokenization: Uses batch_encode_plus with padding to max_length, truncation enabled, and special tokens added. Returns input_ids and attention_mask as numpy arrays.
- Sequence length handling: If max_seq_length is not specified, uses the tokenizer's model_max_length. Includes a guard against a known HuggingFace bug where some models report extremely large max lengths (> 100,000), falling back to max_position_embeddings from the model config.
- Length-based sorting: Optionally sorts rows by token length to improve GPU batching efficiency. Adds a _curator_seq_order column to preserve and later restore original ordering.
Usage
Use TokenizerStage as a preprocessing step before any ModelStage-based inference. It is automatically included when using EmbeddingCreatorStage, but can also be used independently for custom pipelines that need tokenized data.
Code Reference
Source Location
- Repository: NeMo-Curator
- File: nemo_curator/stages/text/models/tokenizer.py
- Lines: 1-170
Signature
class TokenizerStage(ProcessingStage[DocumentBatch, DocumentBatch]):
def __init__(
self,
model_identifier: str,
cache_dir: str | None = None,
hf_token: str | None = None,
text_field: str = "text",
max_chars: int | None = None,
max_seq_length: int | None = None,
padding_side: Literal["left", "right"] = "right",
sort_by_length: bool = True,
unk_token: bool = False,
): ...
Import
from nemo_curator.stages.text.models.tokenizer import TokenizerStage
I/O Contract
Inputs
| Name | Type | Required | Description |
|---|---|---|---|
| model_identifier | str | Yes | HuggingFace model identifier or local path for the tokenizer |
| cache_dir | str or None | No | Directory for caching downloaded model/tokenizer files (default: None) |
| hf_token | str or None | No | HuggingFace authentication token for gated models (default: None) |
| text_field | str | No | Name of the text column in the input DocumentBatch (default: "text") |
| max_chars | int or None | No | Maximum character count for text truncation before tokenization (default: None) |
| max_seq_length | int or None | No | Maximum token sequence length; if None, derived from the model config (default: None) |
| padding_side | Literal["left", "right"] | No | Side on which to pad tokenized sequences (default: "right") |
| sort_by_length | bool | No | Whether to sort rows by token length for GPU batching efficiency (default: True) |
| unk_token | bool | No | If True, sets the pad token to the tokenizer's unknown token (default: False) |
Stage I/O Specification
| Method | Returns |
|---|---|
| inputs() | (["data"], [text_field]) |
| outputs() | (["data"], [text_field, "_curator_input_ids", "_curator_attention_mask"] + optionally ["_curator_seq_order"]) |
Outputs
| Name | Type | Description |
|---|---|---|
| DocumentBatch | DocumentBatch | Input data augmented with _curator_input_ids (list of int per row), _curator_attention_mask (list of int per row), and optionally _curator_seq_order (int) columns |
Usage Examples
Basic Usage
from nemo_curator.stages.text.models.tokenizer import TokenizerStage
tokenizer = TokenizerStage(
model_identifier="sentence-transformers/all-MiniLM-L6-v2",
text_field="text",
max_seq_length=512,
sort_by_length=True,
)
With Character Truncation and Custom Padding
from nemo_curator.stages.text.models.tokenizer import TokenizerStage
tokenizer = TokenizerStage(
model_identifier="intfloat/e5-large-v2",
text_field="content",
max_chars=5000,
padding_side="left",
sort_by_length=False,
)
Using unk_token for Models Without a Pad Token
from nemo_curator.stages.text.models.tokenizer import TokenizerStage
# Some models (e.g., LLaMA) don't define a pad token by default
tokenizer = TokenizerStage(
model_identifier="meta-llama/Llama-2-7b-hf",
hf_token="hf_...",
unk_token=True,
)
Implementation Details
Max Sequence Length Guard
Some HuggingFace models report their model_max_length as an extremely large integer (e.g. int(1e30)). The _setup method detects this case (values greater than 100,000) and falls back to AutoConfig.from_pretrained(...).max_position_embeddings to get the actual maximum position embedding size. The load_cfg method is decorated with @lru_cache(maxsize=1) to avoid redundant config loading.
Sort-by-Length Optimization
When sort_by_length is True, the stage:
- Computes the actual token count per row by summing the attention_mask values
- Adds a _curator_seq_order column with original row indices (0, 1, 2, ...)
- Sorts by token length using stable sort (preserving relative order of equal-length sequences)
- Drops the temporary _curator_token_length column
This sorting ensures that sequences of similar length are batched together during GPU inference, reducing wasted computation on padding tokens. The downstream ModelStage uses the _curator_seq_order column to restore original ordering.
Ray Actor Stage
The ray_stage_spec method returns {"is_actor_stage": True}, indicating that this stage should be scheduled as a Ray actor. This is useful because the tokenizer instance maintains state (the loaded tokenizer model) that should persist across multiple batch invocations.
Tokenization Details
The batch_encode_plus call uses:
- max_length=self.max_seq_length - enforces the maximum sequence length
- padding="max_length" - pads all sequences to the same length
- return_tensors="np" - returns NumPy arrays for efficient conversion
- truncation=True - truncates sequences exceeding max_length
- add_special_tokens=True - adds model-specific special tokens (CLS, SEP, etc.)
- return_token_type_ids=False - omits token type IDs (not needed for most embedding models)
Related Pages
- Implementation:NVIDIA_NeMo_Curator_ModelStage - GPU model inference stage that consumes TokenizerStage output
- Implementation:NVIDIA_NeMo_Curator_EmbedderBase - Composite embedding stage that includes TokenizerStage as its first component
- Implementation:NVIDIA_NeMo_Curator_VLLMEmbedder - Alternative approach that handles tokenization internally
- Environment:NVIDIA_NeMo_Curator_Python_Linux_Base