Implementation:Datajuicer Data juicer InstructionFollowingDifficultyFilter
| Knowledge Sources | |
|---|---|
| Domains | Data_Quality, Filtering |
| Last Updated | 2026-02-14 16:00 GMT |
Overview
Concrete tool for filtering data samples based on their instruction following difficulty (IFD) score provided by Data-Juicer.
Description
InstructionFollowingDifficultyFilter is a filter operator that keeps texts based on their instruction following difficulty (IFD, https://arxiv.org/abs/2308.12032) score. The IFD score is the ratio of the loss with and without the query (loss_w_query / loss_wo_query). The operator computes this using a HuggingFace tokenizer and model. Samples are kept if their IFD score falls within the specified range. It extends LLMPerplexityFilter (which extends Filter) and implements the two-phase compute_stats/process pattern. The IFD score is cached under the ifd_score stats key.
Usage
Import this operator when you need to filter dataset samples based on how difficult they are for a model to follow instructions. Configure it in your Data-Juicer YAML config or instantiate directly.
Code Reference
Source Location
- Repository: Datajuicer_Data_juicer
- File: data_juicer/ops/filter/instruction_following_difficulty_filter.py
- Lines: 1-52
Signature
@OPERATORS.register_module("instruction_following_difficulty_filter")
class InstructionFollowingDifficultyFilter(LLMPerplexityFilter):
# Inherits __init__ from LLMPerplexityFilter
def __init__(
self,
hf_model: str = "Qwen/Qwen2.5-0.5B",
model_params: Optional[Dict] = None,
min_score: float = 1.0,
max_score: float = 100.0,
query_template: Optional[str] = None,
response_template: Optional[str] = None,
*args,
**kwargs,
):
...
Import
from data_juicer.ops.filter.instruction_following_difficulty_filter import InstructionFollowingDifficultyFilter
I/O Contract
Inputs
| Name | Type | Required | Description |
|---|---|---|---|
| hf_model | str | No | HuggingFace model name for computing perplexity/loss. Default: "Qwen/Qwen2.5-0.5B" |
| model_params | Optional[Dict] | No | Parameters for initializing the model. Default: None |
| min_score | float | No | Minimum IFD score to keep samples. Default: 1.0 |
| max_score | float | No | Maximum IFD score to keep samples. Default: 100.0 |
| query_template | Optional[str] | No | Template for building the query string. Default: None |
| response_template | Optional[str] | No | Template for building the response string. Default: None |
Outputs
| Name | Type | Description |
|---|---|---|
| samples | Dict | Filtered samples with stats field updated (ifd_score) |
Usage Examples
YAML Configuration
process:
- instruction_following_difficulty_filter:
hf_model: "Qwen/Qwen2.5-0.5B"
min_score: 1.0
max_score: 100.0
Python API
from data_juicer.ops.filter.instruction_following_difficulty_filter import InstructionFollowingDifficultyFilter
op = InstructionFollowingDifficultyFilter(min_score=1.0, max_score=100.0)
# Apply to dataset
result = dataset.process(op)