Implementation:EvolvingLMMs Lab Lmms eval StructEditBench YAML Config
File: `/tmp/kapso_repo_sslb_59s/lmms_eval/tasks/structeditbench/structeditbench.yaml`
Principle: YAML_Task_Configuration
Overview
The StructEditBench YAML configuration defines an evaluation task for structured visual editing across six categories (chart, math, graph, puzzle, science, table). It uses an OpenAI-compatible VLM judge to score editing quality through question-answering, providing both global metrics and category-specific breakdowns.
Configuration Structure
Dataset Configuration
dataset_path: parquet
dataset_kwargs:
data_files:
train: /path/to/StructEditBench/data/train-*.parquet
task: "structeditbench"
test_split: train
output_type: generate_until
The configuration uses Parquet format with a placeholder path that should be updated for actual usage.
Document Processing
doc_to_visual: !function utils.structeditbench_doc_to_visual
doc_to_text: !function utils.structeditbench_doc_to_text
doc_to_target: !function utils.structeditbench_doc_to_target
All three transformation functions are custom implementations in the utils module.
Generation Parameters
generation_kwargs:
max_new_tokens: 512
temperature: 0
top_p: 1.0
num_beams: 1
do_sample: false
Standard deterministic generation configuration with greedy decoding.
Result Processing
process_results: !function utils.structeditbench_process_results
Expected Dataset Schema
The configuration includes inline documentation of expected fields:
# Expected dataset fields (minimum):
# - source_image (PIL.Image) or input_image/image: source image for editing
# - instruction/edit_prompt/prompt/edit_instruction: edit instruction for the model
# - qa_list: list[{question, ground_truth_answer|answer, label(editing|maintain)}]
# - category: one of {chart, math, graph, puzzle, science, table}
Scoring Backend Configuration
The configuration documents required environment variables for the OpenAI-compatible judge:
# Scoring backend (ImgEdit-style, OpenAI-compatible):
# - Required env vars:
# STRUCTEDITBENCH_API_KEY, STRUCTEDITBENCH_BASE_URL, STRUCTEDITBENCH_EVAL_MODEL_NAME
# - Optional:
# STRUCTEDITBENCH_JUDGE_MODEL_NAME, STRUCTEDITBENCH_TIMEOUT, STRUCTEDITBENCH_MAX_RETRIES, STRUCTEDITBENCH_CALL_DELAY, ...
Metrics
Global Metrics
metric_list:
# Global metrics (percentage, higher is better)
- metric: structeditbench_weighted_accuracy
aggregation: !function utils.structeditbench_aggregate_score
higher_is_better: true
- metric: structeditbench_editing_accuracy
aggregation: !function utils.structeditbench_aggregate_score
higher_is_better: true
- metric: structeditbench_maintain_accuracy
aggregation: !function utils.structeditbench_aggregate_score
higher_is_better: true
Three global metrics track:
- Weighted accuracy: Overall performance with category weighting
- Editing accuracy: Success rate on editing questions
- Maintain accuracy: Success rate on preservation questions
Category-Specific Metrics
# Category breakdown for weighted accuracy
- metric: structeditbench_chart_weighted_accuracy
aggregation: !function utils.structeditbench_aggregate_chart
higher_is_better: true
- metric: structeditbench_math_weighted_accuracy
aggregation: !function utils.structeditbench_aggregate_math
higher_is_better: true
- metric: structeditbench_graph_weighted_accuracy
aggregation: !function utils.structeditbench_aggregate_graph
higher_is_better: true
- metric: structeditbench_puzzle_weighted_accuracy
aggregation: !function utils.structeditbench_aggregate_puzzle
higher_is_better: true
- metric: structeditbench_science_weighted_accuracy
aggregation: !function utils.structeditbench_aggregate_science
higher_is_better: true
- metric: structeditbench_table_weighted_accuracy
aggregation: !function utils.structeditbench_aggregate_table
higher_is_better: true
Each of the six categories has a dedicated weighted accuracy metric.
Prompt Configuration
lmms_eval_specific_kwargs:
default:
pre_prompt: ""
post_prompt: ""
Metadata
metadata:
- version: 0.1
description: "StructEditBench (Structured-Visuals) editing benchmark, scored via QA+judge (OpenAI-compatible VLM)"
Design Patterns
Question-Answering Evaluation
Unlike pixel-based metrics, StructEditBench evaluates edits through question-answering about the edited image, assessing both successful edits and preservation of unrelated content.
External Judge Pattern
The configuration relies on an external OpenAI-compatible VLM to judge answer correctness, allowing for nuanced semantic evaluation rather than exact-match scoring.
Dual-Aspect Evaluation
The benchmark tracks two complementary aspects:
- Editing accuracy: Did the model successfully apply the requested edits?
- Maintain accuracy: Did the model preserve content that should not be changed?
Category Stratification
Performance is tracked across six distinct visual structure types, enabling analysis of model capabilities across different structured visual domains.
Placeholder Path Pattern
The configuration includes a placeholder path (`/path/to/StructEditBench/data/train-*.parquet`) that must be updated for actual usage, serving as clear documentation for users.
Related Components
- Utility functions: `lmms_eval/tasks/structeditbench/utils.py`
- Principle: YAML_Task_Configuration