Implementation:Datajuicer Data juicer ImageCaptioningFromGPT4VMapper
| Knowledge Sources | |
|---|---|
| Domains | Data_Processing, Mapping |
| Last Updated | 2026-02-14 16:00 GMT |
Overview
Concrete tool for generating image captions and QA pairs using GPT-4 Vision API provided by Data-Juicer.
Description
ImageCaptioningFromGPT4VMapper is a mapper operator that generates text captions for images using the GPT-4 Vision model. It encodes images to base64 and sends them to the OpenAI GPT-4V API with a system prompt selected from predefined modes (reasoning, description, conversation, or custom). The generated text can be added to the original sample or replace it depending on keep_original_sample. It operates in batched mode and includes comprehensive error handling for API failures including authentication, rate limiting, network, and timeout errors.
Usage
Use when you need high-quality image descriptions leveraging GPT-4V's advanced visual understanding for creating rich multimodal training datasets with detailed visual reasoning.
Code Reference
Source Location
- Repository: Datajuicer_Data_juicer
- File: data_juicer/ops/mapper/image_captioning_from_gpt4v_mapper.py
Signature
@OPERATORS.register_module("image_captioning_from_gpt4v_mapper")
class ImageCaptioningFromGPT4VMapper(Mapper):
def __init__(self,
mode: str = "description",
api_key: str = "",
max_token: int = 500,
temperature: float = 1.0,
system_prompt: str = "",
user_prompt: str = "",
user_prompt_key: Optional[str] = None,
keep_original_sample: bool = True,
any_or_all: str = "any",
*args, **kwargs):
Import
from data_juicer.ops.mapper.image_captioning_from_gpt4v_mapper import ImageCaptioningFromGPT4VMapper
I/O Contract
Inputs
| Name | Type | Required | Description |
|---|---|---|---|
| mode | str | No | Mode of text generation: reasoning, description, conversation, or custom; defaults to "description" |
| api_key | str | Yes | API key to authenticate the OpenAI request |
| max_token | int | No | Maximum number of tokens to generate, defaults to 500 |
| temperature | float | No | Controls randomness of output (0 to 1), defaults to 1.0 |
| system_prompt | str | No | System prompt for custom mode |
| user_prompt | str | No | User prompt to guide generation for each sample |
| user_prompt_key | Optional[str] | No | Key name in samples to store per-sample prompts; takes precedence over user_prompt |
| keep_original_sample | bool | No | Whether to keep the original sample, defaults to True |
| any_or_all | str | No | Strategy for multi-image samples: any or all, defaults to "any" |
Outputs
| Name | Type | Description |
|---|---|---|
| samples | Dict | Transformed samples with generated captions appended to text field |
Usage Examples
process:
- image_captioning_from_gpt4v_mapper:
mode: "description"
api_key: "your-api-key"
max_token: 500
temperature: 1.0
keep_original_sample: true