Jump to content

Connect SuperML | Leeroopedia MCP: Equip your AI agents with best practices, code verification, and debugging knowledge. Powered by Leeroo — building Organizational Superintelligence. Contact us at founders@leeroo.com.

Implementation:EvolvingLMMs Lab Lmms eval CVRR Utils

From Leeroopedia

File: lmms_eval/tasks/cvrr/utils.py (242 lines)

Principle: Task Utility Functions

Overview

Utility functions for the CVRR-ES (Comprehensive Video Robustness and Reasoning Evaluation - Error Spotting) benchmark. Evaluates video understanding across 11 reasoning dimensions using GPT-based evaluation with correctness and scoring metrics.

Configuration

  • GPT_EVAL_MODEL_NAME - Default: "gpt-4o-2024-11-20"
  • API_TYPE - Default: "openai"
  • NUM_SECONDS_TO_SLEEP - 5 seconds between retries

Evaluation Dimensions

The benchmark covers 11 video understanding dimensions:

  1. Continuity and Object Instance Count
  2. Fine-grained action understanding
  3. Interpretation of social context
  4. Interpretation of visual context
  5. Multiple actions in a single video
  6. Non-existent actions with existent scene depictions
  7. Non-existent actions with non-existent scene depictions
  8. Partial actions
  9. Time order understanding
  10. Understanding of emotional context
  11. Unusual and Physically Anomalous activities

Key Functions

cvrr_doc_to_visual

def cvrr_doc_to_visual(doc)

Locates video file based on dimension category.

Parameters:

  • doc - Document with DimensionName and VideoID fields

Returns: List with video file path

Path Construction:

  1. Uses HF_HOME environment variable
  2. Navigates to CVRR-ES cache directory
  3. Maps DimensionName to subdirectory (snake_case)
  4. Constructs full path: cache_dir/dimension_dir/VideoID
  5. Exits with error if video not found

Dimension Directory Mapping:

  • "Continuity and Object Instance Count" → continuity_and_object_instance_count
  • "Fine-grained action understanding" → fine_grained_action_understanding
  • etc. (all spaces to underscores, lowercase)

cvrr_doc_to_text

def cvrr_doc_to_text(doc, lmms_eval_specific_kwargs=None)

Formats question with optional pre/post prompts.

Parameters:

  • doc - Document with "Q" field
  • lmms_eval_specific_kwargs - Optional pre_prompt and post_prompt

Returns: Formatted question string

cvrr_doc_to_answer

def cvrr_doc_to_answer(doc)

Extracts ground truth answer.

Returns: Answer string from doc["A"]

get_gpt_eval

def get_gpt_eval(question, answer, pred, max_tokens: int, retries: int = 5)

Calls GPT API to evaluate prediction correctness.

Parameters:

  • question - Original question
  • answer - Ground truth answer
  • pred - Model's prediction
  • max_tokens - Maximum response length (512)
  • retries - Retry attempts (default: 5)

Returns: Tuple of (evaluation response, model name)

System Prompt: "You are an intelligent chatbot designed for evaluating the correctness of AI assistant predictions for question-answer pairs..."

Evaluation Instructions:

  • Focus on correctness and accuracy with ground truth
  • Consider predictions with less specific details as correct unless explicitly asked
  • Provide correct/incorrect classification
  • Score on 0 (fully wrong) to 5 (fully correct) integer scale
  • Include reasoning for decision

Output Format: Python dictionary:

{"pred": "correct"|"incorrect", "score": <0-5>, "reason": "<explanation>"}

API Settings:

  • Temperature: 0
  • Timeout: 60 seconds

Error Handling:

  • Catches HTTPError, RequestException, JSONDecodeError
  • Retries with 5-second delays
  • Logs all error types separately
  • Returns empty strings after exhausting retries

parse_score

def parse_score(review)

Parses GPT evaluation response into structured data.

Parameters:

  • review - GPT response string containing dictionary

Returns: Tuple of (correctness, score, reason)

  • correctness - "correct" or "incorrect"
  • score - Integer 0-5
  • reason - String explanation

Error Handling:

  • Catches SyntaxError, ValueError, and general exceptions
  • Logs errors with review content
  • Returns ("incorrect", 0, "") on parsing failures

cvrr_process_results

def cvrr_process_results(doc, result)

Processes single result with GPT evaluation.

Parameters:

  • doc - Document with Q, A, VideoID, DimensionName
  • result - Model's prediction list

Returns: Dictionary with "gpt_eval_score" and "gpt_eval_accuracy" entries

Result Structure:

{
  "gpt_eval_score": {
    "VideoID": ...,
    "Q": ...,
    "A": ...,
    "pred": ...,
    "DimensionName": ...,
    "correctness": "correct"|"incorrect",
    "score": 0-5,
    "reason": "..."
  },
  "gpt_eval_accuracy": { same structure }
}

Fallback: On errors, returns:

  • score: 0
  • correctness: "incorrect"
  • reason: ""
  • Logs error with question_id if available

cvrr_aggregate_score

def cvrr_aggregate_score(results, args)

Calculates average score across all results.

Parameters:

  • results - List of result dictionaries with "score" field
  • args - Additional arguments (unused)

Returns: Average score (0.0-5.0)

Logging: Outputs average score to eval_logger

cvrr_aggregate_accuracy

def cvrr_aggregate_accuracy(results, args)

Calculates accuracy percentage from correctness judgments.

Parameters:

  • results - List of result dictionaries with "correctness" field
  • args - Additional arguments (unused)

Returns: Accuracy as percentage (0-100)

Calculation:

  • yes_count: Number of "correct" judgments
  • no_count: Number of other judgments
  • accuracy = yes_count / (yes_count + no_count) × 100

Logging: Outputs accuracy to eval_logger

Usage Pattern

  1. Map dimension to subdirectory and locate video
  2. Format question with prompts
  3. Get model prediction
  4. Call GPT evaluator (max 512 tokens)
  5. Parse correctness, score (0-5), and reason
  6. Aggregate average score and accuracy percentage
  7. Track performance by dimension

Dependencies

  • ast - Dictionary string parsing
  • os, sys, time - System operations and delays
  • requests - HTTP API calls
  • yaml - Configuration loading
  • loguru - Logging

Page Connections

Double-click a node to navigate. Hold to expand connections.
Principle
Implementation
Heuristic
Environment