Jump to content

Connect SuperML | Leeroopedia MCP: Equip your AI agents with best practices, code verification, and debugging knowledge. Powered by Leeroo — building Organizational Superintelligence. Contact us at founders@leeroo.com.

Implementation:Datajuicer Data juicer InstructionFollowingDifficultyFilter

From Leeroopedia
Knowledge Sources
Domains Data_Quality, Filtering
Last Updated 2026-02-14 16:00 GMT

Overview

Concrete tool for filtering data samples based on their instruction following difficulty (IFD) score provided by Data-Juicer.

Description

InstructionFollowingDifficultyFilter is a filter operator that keeps texts based on their instruction following difficulty (IFD, https://arxiv.org/abs/2308.12032) score. The IFD score is the ratio of the loss with and without the query (loss_w_query / loss_wo_query). The operator computes this using a HuggingFace tokenizer and model. Samples are kept if their IFD score falls within the specified range. It extends LLMPerplexityFilter (which extends Filter) and implements the two-phase compute_stats/process pattern. The IFD score is cached under the ifd_score stats key.

Usage

Import this operator when you need to filter dataset samples based on how difficult they are for a model to follow instructions. Configure it in your Data-Juicer YAML config or instantiate directly.

Code Reference

Source Location

  • Repository: Datajuicer_Data_juicer
  • File: data_juicer/ops/filter/instruction_following_difficulty_filter.py
  • Lines: 1-52

Signature

@OPERATORS.register_module("instruction_following_difficulty_filter")
class InstructionFollowingDifficultyFilter(LLMPerplexityFilter):
    # Inherits __init__ from LLMPerplexityFilter
    def __init__(
        self,
        hf_model: str = "Qwen/Qwen2.5-0.5B",
        model_params: Optional[Dict] = None,
        min_score: float = 1.0,
        max_score: float = 100.0,
        query_template: Optional[str] = None,
        response_template: Optional[str] = None,
        *args,
        **kwargs,
    ):
        ...

Import

from data_juicer.ops.filter.instruction_following_difficulty_filter import InstructionFollowingDifficultyFilter

I/O Contract

Inputs

Name Type Required Description
hf_model str No HuggingFace model name for computing perplexity/loss. Default: "Qwen/Qwen2.5-0.5B"
model_params Optional[Dict] No Parameters for initializing the model. Default: None
min_score float No Minimum IFD score to keep samples. Default: 1.0
max_score float No Maximum IFD score to keep samples. Default: 100.0
query_template Optional[str] No Template for building the query string. Default: None
response_template Optional[str] No Template for building the response string. Default: None

Outputs

Name Type Description
samples Dict Filtered samples with stats field updated (ifd_score)

Usage Examples

YAML Configuration

process:
  - instruction_following_difficulty_filter:
      hf_model: "Qwen/Qwen2.5-0.5B"
      min_score: 1.0
      max_score: 100.0

Python API

from data_juicer.ops.filter.instruction_following_difficulty_filter import InstructionFollowingDifficultyFilter

op = InstructionFollowingDifficultyFilter(min_score=1.0, max_score=100.0)
# Apply to dataset
result = dataset.process(op)

Related Pages

Page Connections

Double-click a node to navigate. Hold to expand connections.
Principle
Implementation
Heuristic
Environment