Implementation:Guardrails ai Guardrails Open Inference: Difference between revisions
Auto-imported from implementations/Guardrails_ai_Guardrails_Open_Inference.md |
Sync from local file |
||
| Line 198: | Line 198: | ||
== Related Pages == | == Related Pages == | ||
* [[Guardrails_ai_Guardrails_Telemetry_Common]] | * [[Implementation:Guardrails_ai_Guardrails_Telemetry_Common]] | ||
* [[Guardrails_ai_Guardrails_Guard_Tracing]] | * [[Implementation:Guardrails_ai_Guardrails_Guard_Tracing]] | ||
* [[Guardrails_ai_Guardrails_Hub_Tracing]] | * [[Implementation:Guardrails_ai_Guardrails_Hub_Tracing]] | ||
[[Category:Implementations]] | [[Category:Implementations]] | ||
Latest revision as of 10:40, 27 September 2026
| Knowledge Sources | |
|---|---|
| Domains | Telemetry, Observability, LLM Integration |
| Last Updated | 2026-02-14 00:00 GMT |
Overview
The Open Inference module provides OpenInference-compatible telemetry functions for tracing generic operations and LLM calls by setting standardized span attributes on the active OpenTelemetry span.
Description
This module implements two primary tracing functions that follow the OpenInference semantic conventions for LLM observability:
trace_operation sets input and output attributes on the current span, recording MIME types and values for both input and output of any operation. Values are serialized to JSON strings using the serialize helper from the telemetry common module.
trace_llm_call provides comprehensive LLM call tracing, setting span attributes for:
- Function calls (JSON-serialized function call details)
- Input messages (role-based message lists with per-message attribute indexing)
- Invocation parameters (model configuration with sensitive values redacted via recursive_key_operation and redact)
- Model name
- Output messages (response message lists)
- Prompt template (template string, variables, and version)
- Token counts (completion, prompt, and total)
Both functions retrieve the current span via get_span and silently return if no span is available. When the OpenInference library is installed, spans are also tagged with OPENINFERENCE_SPAN_KIND = "GUARDRAIL".
Usage
Use trace_operation when you need to record input/output for any generic operation within a traced context. Use trace_llm_call specifically for instrumenting LLM API calls with detailed message, parameter, and token count tracking. These functions are called internally by the guard tracing module during guard execution.
Code Reference
Source Location
- Repository: Guardrails
- File:
guardrails/telemetry/open_inference.py
Signature
def trace_operation(
*,
input_mime_type: Optional[str] = None,
input_value: Optional[Any] = None,
output_mime_type: Optional[str] = None,
output_value: Optional[Any] = None,
)
def trace_llm_call(
*,
function_call: Optional[Dict[str, Any]] = None,
input_messages: Optional[List[Dict[str, Any]]] = None,
invocation_parameters: Optional[Dict[str, Any]] = None,
model_name: Optional[str] = None,
output_messages: Optional[List[Dict[str, Any]]] = None,
prompt_template_template: Optional[str] = None,
prompt_template_variables: Optional[Dict[str, Any]] = None,
prompt_template_version: Optional[str] = None,
token_count_completion: Optional[int] = None,
token_count_prompt: Optional[int] = None,
token_count_total: Optional[int] = None,
)
Import
from guardrails.telemetry.open_inference import trace_operation, trace_llm_call
I/O Contract
trace_operation
| Parameter | Type | Description |
|---|---|---|
input_mime_type |
Optional[str] |
MIME type of the input (e.g., "text/plain", "application/json")
|
input_value |
Optional[Any] |
The input value (serialized to JSON string) |
output_mime_type |
Optional[str] |
MIME type of the output |
output_value |
Optional[Any] |
The output value (serialized to JSON string) |
Span Attributes Set by trace_operation
| Attribute | Type | Description |
|---|---|---|
input.mime_type |
str |
Serialized MIME type of the input |
input.value |
str |
Serialized input value |
output.mime_type |
str |
Serialized MIME type of the output |
output.value |
str |
Serialized output value |
trace_llm_call
| Parameter | Type | Description |
|---|---|---|
function_call |
Optional[Dict[str, Any]] |
Function call details (e.g., {"function_name": "add", "args": [1, 2]})
|
input_messages |
Optional[List[Dict[str, Any]]] |
List of input messages with role and content |
invocation_parameters |
Optional[Dict[str, Any]] |
Model invocation parameters (sensitive values are redacted) |
model_name |
Optional[str] |
The LLM model name (e.g., "gpt-3.5-turbo")
|
output_messages |
Optional[List[Dict[str, Any]]] |
List of output messages from the LLM |
prompt_template_template |
Optional[str] |
The prompt template string |
prompt_template_variables |
Optional[Dict[str, Any]] |
Variables applied to the prompt template |
prompt_template_version |
Optional[str] |
Version of the prompt template |
token_count_completion |
Optional[int] |
Number of tokens in the completion |
token_count_prompt |
Optional[int] |
Number of tokens in the prompt |
token_count_total |
Optional[int] |
Total number of tokens |
Span Attributes Set by trace_llm_call
| Attribute | Type | Description |
|---|---|---|
llm.function_call |
str |
JSON-serialized function call details |
llm.input_messages.{i}.message.{key} |
str |
Per-message input attributes (indexed by message position) |
llm.invocation_parameters |
str |
JSON-serialized invocation parameters with sensitive values redacted |
llm.model_name |
str |
The model name |
llm.output_messages.{i}.message.{key} |
str |
Per-message output attributes (indexed by message position) |
llm.prompt_template.template |
str |
The prompt template string |
llm.prompt_template.variables |
str |
JSON-serialized template variables |
llm.prompt_template.version |
str |
The prompt template version |
llm.token_count.completion |
int |
Completion token count |
llm.token_count.prompt |
int |
Prompt token count |
llm.token_count.total |
int |
Total token count |
Usage Examples
from guardrails.telemetry.open_inference import trace_operation, trace_llm_call
# Trace a generic operation
trace_operation(
input_mime_type="text/plain",
input_value="What is the capital of France?",
output_mime_type="application/json",
output_value={"answer": "Paris"},
)
# Trace an LLM call with full details
trace_llm_call(
model_name="gpt-4",
input_messages=[
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "What is the capital of France?"},
],
output_messages=[
{"role": "assistant", "content": "The capital of France is Paris."},
],
invocation_parameters={
"model": "gpt-4",
"temperature": 0.7,
"api_key": "sk-secret123456", # Will be redacted automatically
},
token_count_prompt=25,
token_count_completion=10,
token_count_total=35,
prompt_template_template="Answer the question: ${question}",
prompt_template_variables={"question": "What is the capital of France?"},
prompt_template_version="v1.0",
)