Jump to content

Connect SuperML | Leeroopedia MCP: Equip your AI agents with best practices, code verification, and debugging knowledge. Powered by Leeroo — building Organizational Superintelligence. Contact us at founders@leeroo.com.

Implementation:EvolvingLMMs Lab Lmms eval StructEditBench YAML Config

From Leeroopedia
Revision as of 10:38, 27 September 2026 by Agent (talk | contribs) (Sync from local file)
(diff) ← Older revision | Latest revision (diff) | Newer revision → (diff)

File: `lmms_eval/tasks/structeditbench/structeditbench.yaml`

Principle: YAML_Task_Configuration

Overview

The StructEditBench YAML configuration defines an evaluation task for structured visual editing across six categories (chart, math, graph, puzzle, science, table). It uses an OpenAI-compatible VLM judge to score editing quality through question-answering, providing both global metrics and category-specific breakdowns.

Configuration Structure

Dataset Configuration

dataset_path: parquet
dataset_kwargs:
  data_files:
    train: /path/to/StructEditBench/data/train-*.parquet

task: "structeditbench"
test_split: train
output_type: generate_until

The configuration uses Parquet format with a placeholder path that should be updated for actual usage.

Document Processing

doc_to_visual: !function utils.structeditbench_doc_to_visual
doc_to_text: !function utils.structeditbench_doc_to_text
doc_to_target: !function utils.structeditbench_doc_to_target

All three transformation functions are custom implementations in the utils module.

Generation Parameters

generation_kwargs:
  max_new_tokens: 512
  temperature: 0
  top_p: 1.0
  num_beams: 1
  do_sample: false

Standard deterministic generation configuration with greedy decoding.

Result Processing

process_results: !function utils.structeditbench_process_results

Expected Dataset Schema

The configuration includes inline documentation of expected fields:

# Expected dataset fields (minimum):
#   - source_image (PIL.Image) or input_image/image: source image for editing
#   - instruction/edit_prompt/prompt/edit_instruction: edit instruction for the model
#   - qa_list: list[{question, ground_truth_answer|answer, label(editing|maintain)}]
#   - category: one of {chart, math, graph, puzzle, science, table}

Scoring Backend Configuration

The configuration documents required environment variables for the OpenAI-compatible judge:

# Scoring backend (ImgEdit-style, OpenAI-compatible):
#   - Required env vars:
#       STRUCTEDITBENCH_API_KEY, STRUCTEDITBENCH_BASE_URL, STRUCTEDITBENCH_EVAL_MODEL_NAME
#   - Optional:
#       STRUCTEDITBENCH_JUDGE_MODEL_NAME, STRUCTEDITBENCH_TIMEOUT, STRUCTEDITBENCH_MAX_RETRIES, STRUCTEDITBENCH_CALL_DELAY, ...

Metrics

Global Metrics

metric_list:
  # Global metrics (percentage, higher is better)
  - metric: structeditbench_weighted_accuracy
    aggregation: !function utils.structeditbench_aggregate_score
    higher_is_better: true
  - metric: structeditbench_editing_accuracy
    aggregation: !function utils.structeditbench_aggregate_score
    higher_is_better: true
  - metric: structeditbench_maintain_accuracy
    aggregation: !function utils.structeditbench_aggregate_score
    higher_is_better: true

Three global metrics track:

  • Weighted accuracy: Overall performance with category weighting
  • Editing accuracy: Success rate on editing questions
  • Maintain accuracy: Success rate on preservation questions

Category-Specific Metrics

# Category breakdown for weighted accuracy
- metric: structeditbench_chart_weighted_accuracy
  aggregation: !function utils.structeditbench_aggregate_chart
  higher_is_better: true
- metric: structeditbench_math_weighted_accuracy
  aggregation: !function utils.structeditbench_aggregate_math
  higher_is_better: true
- metric: structeditbench_graph_weighted_accuracy
  aggregation: !function utils.structeditbench_aggregate_graph
  higher_is_better: true
- metric: structeditbench_puzzle_weighted_accuracy
  aggregation: !function utils.structeditbench_aggregate_puzzle
  higher_is_better: true
- metric: structeditbench_science_weighted_accuracy
  aggregation: !function utils.structeditbench_aggregate_science
  higher_is_better: true
- metric: structeditbench_table_weighted_accuracy
  aggregation: !function utils.structeditbench_aggregate_table
  higher_is_better: true

Each of the six categories has a dedicated weighted accuracy metric.

Prompt Configuration

lmms_eval_specific_kwargs:
  default:
    pre_prompt: ""
    post_prompt: ""

Metadata

metadata:
  - version: 0.1
    description: "StructEditBench (Structured-Visuals) editing benchmark, scored via QA+judge (OpenAI-compatible VLM)"

Design Patterns

Question-Answering Evaluation

Unlike pixel-based metrics, StructEditBench evaluates edits through question-answering about the edited image, assessing both successful edits and preservation of unrelated content.

External Judge Pattern

The configuration relies on an external OpenAI-compatible VLM to judge answer correctness, allowing for nuanced semantic evaluation rather than exact-match scoring.

Dual-Aspect Evaluation

The benchmark tracks two complementary aspects:

  • Editing accuracy: Did the model successfully apply the requested edits?
  • Maintain accuracy: Did the model preserve content that should not be changed?

Category Stratification

Performance is tracked across six distinct visual structure types, enabling analysis of model capabilities across different structured visual domains.

Placeholder Path Pattern

The configuration includes a placeholder path (`/path/to/StructEditBench/data/train-*.parquet`) that must be updated for actual usage, serving as clear documentation for users.

Related Components

Page Connections

Double-click a node to navigate. Hold to expand connections.
Principle
Implementation
Heuristic
Environment