Implementation:Open compass VLMEvalKit SArena CLIP Score
| Field | Value |
|---|---|
| source | VLMEvalKit |
| domain | Vision, Evaluation, Image Generation, CLIP |
Overview
Calculates CLIP-based similarity scores for text-to-image and image-to-image evaluation in the SArena benchmark.
Description
The `CLIPScoreCalculator` class extends `BaseMetric` to compute CLIP scores using the `openai/clip-vit-large-patch14` model. It supports two task types: T2I (text-to-image, comparing generated images against text captions) and I2I (image-to-image, comparing generated images against reference images). Scores are computed in batches using DataLoader for efficiency, with GPU acceleration when available.
Usage
Called internally by the corresponding dataset class during evaluation.
Code Reference
- Source:
vlmeval/dataset/utils/SArena/CLIP_Score.py, Lines: L1-72 - Import:
from vlmeval.dataset.utils.SArena.CLIP_Score import CLIPScoreCalculator
Key Functions:
class CLIPScoreCalculator(BaseMetric):
def CLIP_Score(self, images, captions): ...
def calculate_score(self, batch, batch_size=64, update=True): ...
I/O Contract
| Direction | Description |
|---|---|
| Inputs | A batch dict with 'pred_im' (predicted images) and either 'caption' (T2I) or 'gt_im' (I2I ground truth images) |
| Outputs | Tuple of (average_score, list_of_per_sample_scores) |
Usage Examples
from vlmeval.dataset.utils.SArena.CLIP_Score import CLIPScoreCalculator
calc = CLIPScoreCalculator(task_type='T2I')
avg_score, scores = calc.calculate_score(batch)