Implementation:EvolvingLMMs Lab Lmms eval CVRR Utils
File: lmms_eval/tasks/cvrr/utils.py (242 lines)
Principle: Task Utility Functions
Overview
Utility functions for the CVRR-ES (Comprehensive Video Robustness and Reasoning Evaluation - Error Spotting) benchmark. Evaluates video understanding across 11 reasoning dimensions using GPT-based evaluation with correctness and scoring metrics.
Configuration
GPT_EVAL_MODEL_NAME- Default: "gpt-4o-2024-11-20"API_TYPE- Default: "openai"NUM_SECONDS_TO_SLEEP- 5 seconds between retries
Evaluation Dimensions
The benchmark covers 11 video understanding dimensions:
- Continuity and Object Instance Count
- Fine-grained action understanding
- Interpretation of social context
- Interpretation of visual context
- Multiple actions in a single video
- Non-existent actions with existent scene depictions
- Non-existent actions with non-existent scene depictions
- Partial actions
- Time order understanding
- Understanding of emotional context
- Unusual and Physically Anomalous activities
Key Functions
cvrr_doc_to_visual
def cvrr_doc_to_visual(doc)
Locates video file based on dimension category.
Parameters:
doc- Document with DimensionName and VideoID fields
Returns: List with video file path
Path Construction:
- Uses HF_HOME environment variable
- Navigates to CVRR-ES cache directory
- Maps DimensionName to subdirectory (snake_case)
- Constructs full path: cache_dir/dimension_dir/VideoID
- Exits with error if video not found
Dimension Directory Mapping:
- "Continuity and Object Instance Count" → continuity_and_object_instance_count
- "Fine-grained action understanding" → fine_grained_action_understanding
- etc. (all spaces to underscores, lowercase)
cvrr_doc_to_text
def cvrr_doc_to_text(doc, lmms_eval_specific_kwargs=None)
Formats question with optional pre/post prompts.
Parameters:
doc- Document with "Q" fieldlmms_eval_specific_kwargs- Optional pre_prompt and post_prompt
Returns: Formatted question string
cvrr_doc_to_answer
def cvrr_doc_to_answer(doc)
Extracts ground truth answer.
Returns: Answer string from doc["A"]
get_gpt_eval
def get_gpt_eval(question, answer, pred, max_tokens: int, retries: int = 5)
Calls GPT API to evaluate prediction correctness.
Parameters:
question- Original questionanswer- Ground truth answerpred- Model's predictionmax_tokens- Maximum response length (512)retries- Retry attempts (default: 5)
Returns: Tuple of (evaluation response, model name)
System Prompt: "You are an intelligent chatbot designed for evaluating the correctness of AI assistant predictions for question-answer pairs..."
Evaluation Instructions:
- Focus on correctness and accuracy with ground truth
- Consider predictions with less specific details as correct unless explicitly asked
- Provide correct/incorrect classification
- Score on 0 (fully wrong) to 5 (fully correct) integer scale
- Include reasoning for decision
Output Format: Python dictionary:
{"pred": "correct"|"incorrect", "score": <0-5>, "reason": "<explanation>"}
API Settings:
- Temperature: 0
- Timeout: 60 seconds
Error Handling:
- Catches HTTPError, RequestException, JSONDecodeError
- Retries with 5-second delays
- Logs all error types separately
- Returns empty strings after exhausting retries
parse_score
def parse_score(review)
Parses GPT evaluation response into structured data.
Parameters:
review- GPT response string containing dictionary
Returns: Tuple of (correctness, score, reason)
- correctness - "correct" or "incorrect"
- score - Integer 0-5
- reason - String explanation
Error Handling:
- Catches SyntaxError, ValueError, and general exceptions
- Logs errors with review content
- Returns ("incorrect", 0, "") on parsing failures
cvrr_process_results
def cvrr_process_results(doc, result)
Processes single result with GPT evaluation.
Parameters:
doc- Document with Q, A, VideoID, DimensionNameresult- Model's prediction list
Returns: Dictionary with "gpt_eval_score" and "gpt_eval_accuracy" entries
Result Structure:
{
"gpt_eval_score": {
"VideoID": ...,
"Q": ...,
"A": ...,
"pred": ...,
"DimensionName": ...,
"correctness": "correct"|"incorrect",
"score": 0-5,
"reason": "..."
},
"gpt_eval_accuracy": { same structure }
}
Fallback: On errors, returns:
- score: 0
- correctness: "incorrect"
- reason: ""
- Logs error with question_id if available
cvrr_aggregate_score
def cvrr_aggregate_score(results, args)
Calculates average score across all results.
Parameters:
results- List of result dictionaries with "score" fieldargs- Additional arguments (unused)
Returns: Average score (0.0-5.0)
Logging: Outputs average score to eval_logger
cvrr_aggregate_accuracy
def cvrr_aggregate_accuracy(results, args)
Calculates accuracy percentage from correctness judgments.
Parameters:
results- List of result dictionaries with "correctness" fieldargs- Additional arguments (unused)
Returns: Accuracy as percentage (0-100)
Calculation:
- yes_count: Number of "correct" judgments
- no_count: Number of other judgments
- accuracy = yes_count / (yes_count + no_count) × 100
Logging: Outputs accuracy to eval_logger
Usage Pattern
- Map dimension to subdirectory and locate video
- Format question with prompts
- Get model prediction
- Call GPT evaluator (max 512 tokens)
- Parse correctness, score (0-5), and reason
- Aggregate average score and accuracy percentage
- Track performance by dimension
Dependencies
ast- Dictionary string parsingos, sys, time- System operations and delaysrequests- HTTP API callsyaml- Configuration loadingloguru- Logging