Jump to content

Connect SuperML | Leeroopedia MCP: Equip your AI agents with best practices, code verification, and debugging knowledge. Powered by Leeroo — building Organizational Superintelligence. Contact us at founders@leeroo.com.

Implementation:Hiyouga LLaMA Factory Data Args

From Leeroopedia


Knowledge Sources
Domains Configuration, Data Processing
Last Updated 2026-02-06 19:00 GMT

Overview

Dataclass defining all dataset and data processing configuration arguments for training and evaluation workflows.

Description

The DataArguments dataclass specifies fields for template selection, dataset names and paths, media directories, tokenization cutoff length, streaming mode, sequence packing (including neat packing), dataset mixing strategies (concat, interleave with under/over/once sampling), preprocessing parallelism, validation split configuration, and various training behavior flags. The __post_init__ method performs validation: it splits comma-separated dataset names, resolves the media directory default, checks for incompatible option combinations (e.g., streaming with max_samples, mask_history with train_on_prompt), validates interleave probability lengths against dataset counts, enables packing when neat_packing is set, and decrements cutoff_len by 1 when packing is active to avoid pad_to_multiple_of issues. The to_dict method provides serialization to a plain dictionary.

Usage

Use this dataclass as part of the argument parsing pipeline when configuring a training or evaluation run. It is typically instantiated via HuggingFace's HfArgumentParser alongside model and training arguments. All data processors and dataset loaders read their configuration from an instance of DataArguments.

Code Reference

Source Location

Signature

@dataclass
class DataArguments:
    template: str | None = None
    dataset: str | None = None
    eval_dataset: str | None = None
    dataset_dir: str = "data"
    media_dir: str | None = None
    cutoff_len: int = 2048
    train_on_prompt: bool = False
    mask_history: bool = False
    streaming: bool = False
    buffer_size: int = 16384
    mix_strategy: Literal["concat", "interleave_under", "interleave_over", "interleave_once"] = "concat"
    interleave_probs: str | None = None
    overwrite_cache: bool = False
    preprocessing_batch_size: int = 1000
    preprocessing_num_workers: int | None = None
    max_samples: int | None = None
    eval_num_beams: int | None = None
    ignore_pad_token_for_loss: bool = True
    val_size: float = 0.0
    eval_on_each_dataset: bool = False
    packing: bool | None = None
    neat_packing: bool = False
    tool_format: str | None = None
    default_system: str | None = None
    enable_thinking: bool | None = True
    tokenized_path: str | None = None
    data_shared_file_system: bool = False

    def __post_init__(self) -> None
    def to_dict(self) -> dict[str, Any]

Import

from llamafactory.hparams.data_args import DataArguments

I/O Contract

Inputs

Name Type Required Description
template None No Template name for prompt construction
dataset None No Comma-separated dataset names for training
eval_dataset None No Comma-separated dataset names for evaluation
dataset_dir str No Path to dataset configuration directory (default: "data")
cutoff_len int No Maximum tokenized sequence length (default: 2048)
train_on_prompt bool No Whether to include prompt tokens in loss computation (default: False)
mask_history bool No Whether to mask conversation history and train only on the last turn (default: False)
streaming bool No Enable dataset streaming mode (default: False)
packing None No Enable sequence packing; auto-enabled for pre-training (default: None)
neat_packing bool No Enable packing without cross-attention between packed examples (default: False)
mix_strategy str No Dataset mixing strategy: concat, interleave_under, interleave_over, interleave_once (default: "concat")

Outputs

Name Type Description
DataArguments instance DataArguments Validated configuration object with parsed dataset lists and adjusted parameters
to_dict() dict[str, Any] Dictionary representation of all arguments

Usage Examples

from llamafactory.hparams.data_args import DataArguments

# Create with defaults
data_args = DataArguments(
    template="llama3",
    dataset="alpaca_en,alpaca_zh",
    dataset_dir="data",
    cutoff_len=4096,
    packing=True,
    neat_packing=True,
)
# After __post_init__:
# data_args.dataset == ["alpaca_en", "alpaca_zh"]
# data_args.packing == True
# data_args.cutoff_len == 4095  (decremented by 1 for packing)
# Use with HfArgumentParser
from transformers import HfArgumentParser
from llamafactory.hparams.data_args import DataArguments

parser = HfArgumentParser(DataArguments)
data_args = parser.parse_args_into_dataclasses()[0]
print(data_args.to_dict())

Related Pages

Page Connections

Double-click a node to navigate. Hold to expand connections.
Principle
Implementation
Heuristic
Environment