Jump to content

Connect SuperML | Leeroopedia MCP: Equip your AI agents with best practices, code verification, and debugging knowledge. Powered by Leeroo — building Organizational Superintelligence. Contact us at founders@leeroo.com.

Implementation:Datajuicer Data juicer OptimizeResponseMapper

From Leeroopedia
Revision as of 12:22, 16 February 2026 by Admin (talk | contribs) (Auto-imported from implementations/Datajuicer_Data_juicer_OptimizeResponseMapper.md)
(diff) ← Older revision | Latest revision (diff) | Newer revision → (diff)
Knowledge Sources
Domains Data_Processing, Mapping
Last Updated 2026-02-14 16:00 GMT

Overview

Concrete tool for optimizing responses in question-answer pairs provided by Data-Juicer.

Description

OptimizeResponseMapper is a mapper operator that optimizes only the response (answer) in a QA pair to be more detailed and specific, while ensuring it still addresses the original question. It extends OptimizeQAMapper with a specialized Chinese system prompt and overrides parse_output to return None for the query and the stripped raw output as the optimized response, leaving the original question unchanged. Requires CUDA acceleration.

Usage

Use when you have well-formed questions but need to enhance answer quality to produce more comprehensive or detailed responses for instruction-following training data.

Code Reference

Source Location

Signature

@OPERATORS.register_module("optimize_response_mapper")
class OptimizeResponseMapper(OptimizeQAMapper):
    def __init__(
        self,
        api_or_hf_model: str = "Qwen/Qwen2.5-7B-Instruct",
        is_hf_model: bool = True,
        *,
        api_endpoint: Optional[str] = None,
        response_path: Optional[str] = None,
        system_prompt: Optional[str] = None,
        input_template: Optional[str] = None,
        qa_pair_template: Optional[str] = None,
        output_pattern: Optional[str] = None,
        try_num: PositiveInt = 3,
        enable_vllm: bool = False,
        model_params: Optional[Dict] = None,
        sampling_params: Optional[Dict] = None,
        **kwargs,
    ):

Import

from data_juicer.ops.mapper.optimize_response_mapper import OptimizeResponseMapper

I/O Contract

Inputs

Name Type Required Description
api_or_hf_model str No API or HuggingFace model name (default: Qwen/Qwen2.5-7B-Instruct)
is_hf_model bool No Whether to use HuggingFace model (default: True)
enable_vllm bool No Whether to use VLLM for inference acceleration (default: False)
try_num PositiveInt No Number of retry attempts on error (default: 3)
sample[query_key] str Yes The original question in the QA pair
sample[response_key] str Yes The response to optimize

Outputs

Name Type Description
sample[response_key] str Optimized, more detailed response

Usage Examples

process:
  - optimize_response_mapper:
      api_or_hf_model: Qwen/Qwen2.5-7B-Instruct
      is_hf_model: true
      enable_vllm: true

Related Pages

Page Connections

Double-click a node to navigate. Hold to expand connections.
Principle
Implementation
Heuristic
Environment