Jump to content

Connect SuperML | Leeroopedia MCP: Equip your AI agents with best practices, code verification, and debugging knowledge. Powered by Leeroo — building Organizational Superintelligence. Contact us at founders@leeroo.com.

Implementation:Datajuicer Data juicer ImageCaptioningFromGPT4VMapper

From Leeroopedia
Knowledge Sources
Domains Data_Processing, Mapping
Last Updated 2026-02-14 16:00 GMT

Overview

Concrete tool for generating image captions and QA pairs using GPT-4 Vision API provided by Data-Juicer.

Description

ImageCaptioningFromGPT4VMapper is a mapper operator that generates text captions for images using the GPT-4 Vision model. It encodes images to base64 and sends them to the OpenAI GPT-4V API with a system prompt selected from predefined modes (reasoning, description, conversation, or custom). The generated text can be added to the original sample or replace it depending on keep_original_sample. It operates in batched mode and includes comprehensive error handling for API failures including authentication, rate limiting, network, and timeout errors.

Usage

Use when you need high-quality image descriptions leveraging GPT-4V's advanced visual understanding for creating rich multimodal training datasets with detailed visual reasoning.

Code Reference

Source Location

Signature

@OPERATORS.register_module("image_captioning_from_gpt4v_mapper")
class ImageCaptioningFromGPT4VMapper(Mapper):
    def __init__(self,
                 mode: str = "description",
                 api_key: str = "",
                 max_token: int = 500,
                 temperature: float = 1.0,
                 system_prompt: str = "",
                 user_prompt: str = "",
                 user_prompt_key: Optional[str] = None,
                 keep_original_sample: bool = True,
                 any_or_all: str = "any",
                 *args, **kwargs):

Import

from data_juicer.ops.mapper.image_captioning_from_gpt4v_mapper import ImageCaptioningFromGPT4VMapper

I/O Contract

Inputs

Name Type Required Description
mode str No Mode of text generation: reasoning, description, conversation, or custom; defaults to "description"
api_key str Yes API key to authenticate the OpenAI request
max_token int No Maximum number of tokens to generate, defaults to 500
temperature float No Controls randomness of output (0 to 1), defaults to 1.0
system_prompt str No System prompt for custom mode
user_prompt str No User prompt to guide generation for each sample
user_prompt_key Optional[str] No Key name in samples to store per-sample prompts; takes precedence over user_prompt
keep_original_sample bool No Whether to keep the original sample, defaults to True
any_or_all str No Strategy for multi-image samples: any or all, defaults to "any"

Outputs

Name Type Description
samples Dict Transformed samples with generated captions appended to text field

Usage Examples

process:
  - image_captioning_from_gpt4v_mapper:
      mode: "description"
      api_key: "your-api-key"
      max_token: 500
      temperature: 1.0
      keep_original_sample: true

Related Pages

Page Connections

Double-click a node to navigate. Hold to expand connections.
Principle
Implementation
Heuristic
Environment