Implementation:Open compass VLMEvalKit MMVet Utils
| Field | Value |
|---|---|
| source | VLMEvalKit |
| domain | Vision, Evaluation, LLM Judge, Multi-capability |
Overview
Provides GPT-based scoring utilities for the MM-Vet benchmark, evaluating model predictions against ground truth with AND/OR logic support.
Description
This module implements `build_mmvet_gpt4_prompt` which constructs scoring prompts with few-shot examples demonstrating the AND/OR evaluation logic: AND requires all ground truth elements present for full credit, OR requires any one element. The `MMVet_auxeval` function uses a retry mechanism (up to 5 attempts) to obtain a float score (0.0 to 1.0 in 0.1 increments) from a GPT judge, with increasing temperature for robustness.
Usage
Called internally by the corresponding dataset class during evaluation.
Code Reference
- Source:
vlmeval/dataset/utils/mmvet.py, Lines: L1-106 - Import:
from vlmeval.dataset.utils.mmvet import MMVet_auxeval, build_mmvet_gpt4_prompt
Key Functions:
def build_mmvet_gpt4_prompt(line): ...
def MMVet_auxeval(model, line): ...
I/O Contract
| Direction | Description |
|---|---|
| Inputs | A data line dict with 'question', 'answer' (with AND/OR operators), and 'prediction' fields; a judge model |
| Outputs | Dict with 'log' (evaluation trace) and 'res' (float score 0.0-1.0) |
Usage Examples
from vlmeval.dataset.utils.mmvet import MMVet_auxeval
result = MMVet_auxeval(judge_model, line)