Implementation:Vibrantlabsai Ragas NVMetrics
| Knowledge Sources | |
|---|---|
| Domains | LLM Evaluation, RAG Metrics, NVIDIA NIM |
| Last Updated | 2026-02-12 00:00 GMT |
Overview
The NVMetrics module provides three NVIDIA-developed LLM-as-a-Judge evaluation metrics -- AnswerAccuracy, ContextRelevance, and ResponseGroundedness -- that use dual prompt templates with score averaging for robust assessment of RAG pipeline quality.
Description
This module implements three evaluation metrics designed by NVIDIA for assessing RAG (Retrieval-Augmented Generation) systems. All three metrics share a common architectural pattern: they use two distinct prompt templates per evaluation, run them independently, and average the two scores for robustness. Each metric also includes a retry mechanism (default 5 retries) that re-queries the LLM when it fails to produce a parseable score. NaN detection uses the identity check score == score (NaN is the only float that does not equal itself) to detect parsing failures.
AnswerAccuracy evaluates whether the LLM's response matches a reference answer given a question. It uses two prompt templates (template_accuracy1 and template_accuracy2) that frame the rating task differently -- one as an instruction to rate, the other as a self-description of the rating process. Both use a 3-point scale (0, 2, 4) normalized to [0, 1]. The metric was benchmarked on a leaderboard of zero-shot LLM judges, with nvidia/Llama-3_3-Nemotron-Super-49B-v1 achieving the highest correlation with human judges (~0.92). Required columns: user_input, response, reference.
ContextRelevance scores how relevant retrieved contexts are to the user's question. It uses two prompt templates (template_relevance1 and template_relevance2) with a 3-point scale (0, 1, 2) normalized to [0, 1]. It includes early-return guards for empty inputs, identical question-context pairs, or contexts contained within the question. Required columns: user_input, retrieved_contexts.
ResponseGroundedness evaluates whether the response is grounded in (supported by) the retrieved contexts. It uses two prompt templates (template_groundedness1 and template_groundedness2) with a 3-point scale (0, 1, 2) normalized to [0, 1]. It includes early-return guards for empty inputs (returns 0.0), identical response-context pairs, or response contained in contexts (returns 1.0). Required columns: response, retrieved_contexts.
All metrics use a low temperature (0.1) for LLM generation to ensure deterministic scoring. The process_score method extracts the numeric rating from the LLM's text response, and average_scores combines the two template scores, falling back to the maximum if one is NaN.
Usage
Import these metrics when evaluating RAG pipelines with NVIDIA-style prompt-based scoring. They are designed to work with any Ragas-compatible LLM but were specifically benchmarked with NVIDIA NIM models. Use AnswerAccuracy for answer quality against references, ContextRelevance for retrieval relevance, and ResponseGroundedness for response faithfulness to retrieved contexts.
Code Reference
Source Location
- Repository: Vibrantlabsai_Ragas
- File: src/ragas/metrics/_nv_metrics.py
Signature
@dataclass
class AnswerAccuracy(MetricWithLLM, SingleTurnMetric):
name: str = "nv_accuracy"
template_accuracy1: str = ... # Instruction-style prompt
template_accuracy2: str = ... # Self-description-style prompt
retry: int = 5
def process_score(self, response) -> float: ...
def average_scores(self, score0, score1) -> float: ...
async def _single_turn_ascore(
self, sample: SingleTurnSample, callbacks: Callbacks
) -> float: ...
@dataclass
class ContextRelevance(MetricWithLLM, SingleTurnMetric):
name: str = "nv_context_relevance"
template_relevance1: str = ...
template_relevance2: str = ...
retry: int = 5
def process_score(self, response) -> float: ...
def average_scores(self, score0, score1) -> float: ...
async def _single_turn_ascore(
self, sample: SingleTurnSample, callbacks: Callbacks
) -> float: ...
@dataclass
class ResponseGroundedness(MetricWithLLM, SingleTurnMetric):
name: str = "nv_response_groundedness"
template_groundedness1: str = ...
template_groundedness2: str = ...
retry: int = 5
def process_score(self, response) -> float: ...
def average_scores(self, score0, score1) -> float: ...
async def _single_turn_ascore(
self, sample: SingleTurnSample, callbacks: Callbacks
) -> float: ...
Import
from ragas.metrics._nv_metrics import AnswerAccuracy, ContextRelevance, ResponseGroundedness
I/O Contract
Inputs (AnswerAccuracy)
| Name | Type | Required | Description |
|---|---|---|---|
| user_input | str | Yes | The user's question |
| response | str | Yes | The LLM-generated response to evaluate |
| reference | str | Yes | The ground truth reference answer |
Inputs (ContextRelevance)
| Name | Type | Required | Description |
|---|---|---|---|
| user_input | str | Yes | The user's question |
| retrieved_contexts | List[str] | Yes | Retrieved context strings to evaluate for relevance |
Inputs (ResponseGroundedness)
| Name | Type | Required | Description |
|---|---|---|---|
| response | str | Yes | The LLM-generated response to evaluate |
| retrieved_contexts | List[str] | Yes | Retrieved context strings that should ground the response |
Outputs
| Name | Type | Description |
|---|---|---|
| score (AnswerAccuracy) | float | Normalized score in [0, 0.25, 0.5, 0.75, 1.0] representing accuracy; NaN on failure |
| score (ContextRelevance) | float | Normalized score in [0, 0.25, 0.5, 0.75, 1.0] representing relevance; NaN on failure |
| score (ResponseGroundedness) | float | Normalized score in [0, 0.25, 0.5, 0.75, 1.0] representing groundedness; NaN on failure |
Usage Examples
AnswerAccuracy
from ragas.metrics._nv_metrics import AnswerAccuracy
from ragas.dataset_schema import SingleTurnSample
from ragas.llms import llm_factory
llm = llm_factory("gpt-4o-mini")
metric = AnswerAccuracy()
metric.llm = llm
sample = SingleTurnSample(
user_input="What is the capital of France?",
response="The capital of France is Paris.",
reference="Paris is the capital of France.",
)
score = await metric.single_turn_ascore(sample)
print(f"Answer Accuracy: {score}")
ContextRelevance
from ragas.metrics._nv_metrics import ContextRelevance
from ragas.dataset_schema import SingleTurnSample
metric = ContextRelevance()
metric.llm = llm
sample = SingleTurnSample(
user_input="What is quantum computing?",
retrieved_contexts=[
"Quantum computing uses quantum bits or qubits that can exist in superposition states.",
"Classical computers use binary bits that are either 0 or 1.",
],
)
score = await metric.single_turn_ascore(sample)
print(f"Context Relevance: {score}")
ResponseGroundedness
from ragas.metrics._nv_metrics import ResponseGroundedness
from ragas.dataset_schema import SingleTurnSample
metric = ResponseGroundedness()
metric.llm = llm
sample = SingleTurnSample(
response="Quantum computers use qubits that leverage superposition.",
retrieved_contexts=[
"Quantum computing uses quantum bits or qubits that can exist in superposition states.",
],
)
score = await metric.single_turn_ascore(sample)
print(f"Response Groundedness: {score}")
Scoring Methodology
Dual Prompt Architecture
Each metric uses two prompt templates that frame the evaluation task differently. This dual-prompt design reduces bias from any single prompt formulation. The two scores are averaged:
def average_scores(self, score0, score1):
if score0 >= 0 and score1 >= 0:
score = (score0 + score1) / 2
else:
score = max(score0, score1)
return score
If both scores are valid (non-NaN, non-negative), they are averaged. If one fails, the other is used as fallback.
Score Scales
| Metric | Raw Scale | Normalization | Possible Normalized Values |
|---|---|---|---|
| AnswerAccuracy | 0, 2, 4 | Divide by 4 | 0.0, 0.5, 1.0 |
| ContextRelevance | 0, 1, 2 | Divide by 2 | 0.0, 0.5, 1.0 |
| ResponseGroundedness | 0, 1, 2 | Divide by 2 | 0.0, 0.5, 1.0 |
Since two scores are averaged, the final score can take intermediate values (e.g., 0.25, 0.75).
NaN Handling
The retry loop uses the idiom score == score to detect NaN (since NaN != NaN in IEEE 754 floating point). If the LLM fails to produce a parseable score after all retries, np.nan is returned. If an exception occurs during the entire scoring process, np.nan is assigned and the sample is skipped.
Edge Case Guards
ContextRelevance returns 0.0 early when:
- The user input or context is empty
- The user input equals the context
- The context is contained within the user input
ResponseGroundedness returns 0.0 when response or context is empty, and 1.0 when:
- The response equals the context
- The response is contained within the context
Related Pages
- MetricWithLLM - Mixin providing LLM integration for metrics
- SingleTurnMetric - Base class for single-turn evaluation metrics
- BaseRagasLLM - The LLM interface used for text generation via agenerate_text
- Vibrantlabsai_Ragas_ContextPrecision - Context precision metric for retrieval ranking quality
- Vibrantlabsai_Ragas_FactualCorrectness - Factual correctness metric using claim decomposition and NLI