Jump to content

Connect SuperML | Leeroopedia MCP: Equip your AI agents with best practices, code verification, and debugging knowledge. Powered by Leeroo — building Organizational Superintelligence. Contact us at founders@leeroo.com.

Implementation:Vibrantlabsai Ragas NVMetrics

From Leeroopedia
Knowledge Sources
Domains LLM Evaluation, RAG Metrics, NVIDIA NIM
Last Updated 2026-02-12 00:00 GMT

Overview

The NVMetrics module provides three NVIDIA-developed LLM-as-a-Judge evaluation metrics -- AnswerAccuracy, ContextRelevance, and ResponseGroundedness -- that use dual prompt templates with score averaging for robust assessment of RAG pipeline quality.

Description

This module implements three evaluation metrics designed by NVIDIA for assessing RAG (Retrieval-Augmented Generation) systems. All three metrics share a common architectural pattern: they use two distinct prompt templates per evaluation, run them independently, and average the two scores for robustness. Each metric also includes a retry mechanism (default 5 retries) that re-queries the LLM when it fails to produce a parseable score. NaN detection uses the identity check score == score (NaN is the only float that does not equal itself) to detect parsing failures.

AnswerAccuracy evaluates whether the LLM's response matches a reference answer given a question. It uses two prompt templates (template_accuracy1 and template_accuracy2) that frame the rating task differently -- one as an instruction to rate, the other as a self-description of the rating process. Both use a 3-point scale (0, 2, 4) normalized to [0, 1]. The metric was benchmarked on a leaderboard of zero-shot LLM judges, with nvidia/Llama-3_3-Nemotron-Super-49B-v1 achieving the highest correlation with human judges (~0.92). Required columns: user_input, response, reference.

ContextRelevance scores how relevant retrieved contexts are to the user's question. It uses two prompt templates (template_relevance1 and template_relevance2) with a 3-point scale (0, 1, 2) normalized to [0, 1]. It includes early-return guards for empty inputs, identical question-context pairs, or contexts contained within the question. Required columns: user_input, retrieved_contexts.

ResponseGroundedness evaluates whether the response is grounded in (supported by) the retrieved contexts. It uses two prompt templates (template_groundedness1 and template_groundedness2) with a 3-point scale (0, 1, 2) normalized to [0, 1]. It includes early-return guards for empty inputs (returns 0.0), identical response-context pairs, or response contained in contexts (returns 1.0). Required columns: response, retrieved_contexts.

All metrics use a low temperature (0.1) for LLM generation to ensure deterministic scoring. The process_score method extracts the numeric rating from the LLM's text response, and average_scores combines the two template scores, falling back to the maximum if one is NaN.

Usage

Import these metrics when evaluating RAG pipelines with NVIDIA-style prompt-based scoring. They are designed to work with any Ragas-compatible LLM but were specifically benchmarked with NVIDIA NIM models. Use AnswerAccuracy for answer quality against references, ContextRelevance for retrieval relevance, and ResponseGroundedness for response faithfulness to retrieved contexts.

Code Reference

Source Location

Signature

@dataclass
class AnswerAccuracy(MetricWithLLM, SingleTurnMetric):
    name: str = "nv_accuracy"
    template_accuracy1: str = ...  # Instruction-style prompt
    template_accuracy2: str = ...  # Self-description-style prompt
    retry: int = 5

    def process_score(self, response) -> float: ...
    def average_scores(self, score0, score1) -> float: ...
    async def _single_turn_ascore(
        self, sample: SingleTurnSample, callbacks: Callbacks
    ) -> float: ...


@dataclass
class ContextRelevance(MetricWithLLM, SingleTurnMetric):
    name: str = "nv_context_relevance"
    template_relevance1: str = ...
    template_relevance2: str = ...
    retry: int = 5

    def process_score(self, response) -> float: ...
    def average_scores(self, score0, score1) -> float: ...
    async def _single_turn_ascore(
        self, sample: SingleTurnSample, callbacks: Callbacks
    ) -> float: ...


@dataclass
class ResponseGroundedness(MetricWithLLM, SingleTurnMetric):
    name: str = "nv_response_groundedness"
    template_groundedness1: str = ...
    template_groundedness2: str = ...
    retry: int = 5

    def process_score(self, response) -> float: ...
    def average_scores(self, score0, score1) -> float: ...
    async def _single_turn_ascore(
        self, sample: SingleTurnSample, callbacks: Callbacks
    ) -> float: ...

Import

from ragas.metrics._nv_metrics import AnswerAccuracy, ContextRelevance, ResponseGroundedness

I/O Contract

Inputs (AnswerAccuracy)

Name Type Required Description
user_input str Yes The user's question
response str Yes The LLM-generated response to evaluate
reference str Yes The ground truth reference answer

Inputs (ContextRelevance)

Name Type Required Description
user_input str Yes The user's question
retrieved_contexts List[str] Yes Retrieved context strings to evaluate for relevance

Inputs (ResponseGroundedness)

Name Type Required Description
response str Yes The LLM-generated response to evaluate
retrieved_contexts List[str] Yes Retrieved context strings that should ground the response

Outputs

Name Type Description
score (AnswerAccuracy) float Normalized score in [0, 0.25, 0.5, 0.75, 1.0] representing accuracy; NaN on failure
score (ContextRelevance) float Normalized score in [0, 0.25, 0.5, 0.75, 1.0] representing relevance; NaN on failure
score (ResponseGroundedness) float Normalized score in [0, 0.25, 0.5, 0.75, 1.0] representing groundedness; NaN on failure

Usage Examples

AnswerAccuracy

from ragas.metrics._nv_metrics import AnswerAccuracy
from ragas.dataset_schema import SingleTurnSample
from ragas.llms import llm_factory

llm = llm_factory("gpt-4o-mini")
metric = AnswerAccuracy()
metric.llm = llm

sample = SingleTurnSample(
    user_input="What is the capital of France?",
    response="The capital of France is Paris.",
    reference="Paris is the capital of France.",
)

score = await metric.single_turn_ascore(sample)
print(f"Answer Accuracy: {score}")

ContextRelevance

from ragas.metrics._nv_metrics import ContextRelevance
from ragas.dataset_schema import SingleTurnSample

metric = ContextRelevance()
metric.llm = llm

sample = SingleTurnSample(
    user_input="What is quantum computing?",
    retrieved_contexts=[
        "Quantum computing uses quantum bits or qubits that can exist in superposition states.",
        "Classical computers use binary bits that are either 0 or 1.",
    ],
)

score = await metric.single_turn_ascore(sample)
print(f"Context Relevance: {score}")

ResponseGroundedness

from ragas.metrics._nv_metrics import ResponseGroundedness
from ragas.dataset_schema import SingleTurnSample

metric = ResponseGroundedness()
metric.llm = llm

sample = SingleTurnSample(
    response="Quantum computers use qubits that leverage superposition.",
    retrieved_contexts=[
        "Quantum computing uses quantum bits or qubits that can exist in superposition states.",
    ],
)

score = await metric.single_turn_ascore(sample)
print(f"Response Groundedness: {score}")

Scoring Methodology

Dual Prompt Architecture

Each metric uses two prompt templates that frame the evaluation task differently. This dual-prompt design reduces bias from any single prompt formulation. The two scores are averaged:

def average_scores(self, score0, score1):
    if score0 >= 0 and score1 >= 0:
        score = (score0 + score1) / 2
    else:
        score = max(score0, score1)
    return score

If both scores are valid (non-NaN, non-negative), they are averaged. If one fails, the other is used as fallback.

Score Scales

Metric Raw Scale Normalization Possible Normalized Values
AnswerAccuracy 0, 2, 4 Divide by 4 0.0, 0.5, 1.0
ContextRelevance 0, 1, 2 Divide by 2 0.0, 0.5, 1.0
ResponseGroundedness 0, 1, 2 Divide by 2 0.0, 0.5, 1.0

Since two scores are averaged, the final score can take intermediate values (e.g., 0.25, 0.75).

NaN Handling

The retry loop uses the idiom score == score to detect NaN (since NaN != NaN in IEEE 754 floating point). If the LLM fails to produce a parseable score after all retries, np.nan is returned. If an exception occurs during the entire scoring process, np.nan is assigned and the sample is skipped.

Edge Case Guards

ContextRelevance returns 0.0 early when:

  • The user input or context is empty
  • The user input equals the context
  • The context is contained within the user input

ResponseGroundedness returns 0.0 when response or context is empty, and 1.0 when:

  • The response equals the context
  • The response is contained within the context

Related Pages

Page Connections

Double-click a node to navigate. Hold to expand connections.
Principle
Implementation
Heuristic
Environment