Jump to content

Connect SuperML | Leeroopedia MCP: Equip your AI agents with best practices, code verification, and debugging knowledge. Powered by Leeroo — building Organizational Superintelligence. Contact us at founders@leeroo.com.

Implementation:Open compass VLMEvalKit MMVet Utils

From Leeroopedia
Field Value
source VLMEvalKit
domain Vision, Evaluation, LLM Judge, Multi-capability

Overview

Provides GPT-based scoring utilities for the MM-Vet benchmark, evaluating model predictions against ground truth with AND/OR logic support.

Description

This module implements `build_mmvet_gpt4_prompt` which constructs scoring prompts with few-shot examples demonstrating the AND/OR evaluation logic: AND requires all ground truth elements present for full credit, OR requires any one element. The `MMVet_auxeval` function uses a retry mechanism (up to 5 attempts) to obtain a float score (0.0 to 1.0 in 0.1 increments) from a GPT judge, with increasing temperature for robustness.

Usage

Called internally by the corresponding dataset class during evaluation.

Code Reference

  • Source: vlmeval/dataset/utils/mmvet.py, Lines: L1-106
  • Import: from vlmeval.dataset.utils.mmvet import MMVet_auxeval, build_mmvet_gpt4_prompt

Key Functions:

def build_mmvet_gpt4_prompt(line): ...
def MMVet_auxeval(model, line): ...

I/O Contract

Direction Description
Inputs A data line dict with 'question', 'answer' (with AND/OR operators), and 'prediction' fields; a judge model
Outputs Dict with 'log' (evaluation trace) and 'res' (float score 0.0-1.0)

Usage Examples

from vlmeval.dataset.utils.mmvet import MMVet_auxeval

result = MMVet_auxeval(judge_model, line)

Related Pages

Page Connections

Double-click a node to navigate. Hold to expand connections.
Principle
Implementation
Heuristic
Environment