Implementation:Run llama Llama index BaseVoiceAgent: Difference between revisions
Auto-imported from implementations/Run_llama_Llama_index_BaseVoiceAgent.md |
Sync from local file |
||
| Line 218: | Line 218: | ||
== See Also == | == See Also == | ||
* [[Run_llama_Llama_index_BaseVoiceAgentWebsocket]] - Websocket base class used by voice agents | * [[Implementation:Run_llama_Llama_index_BaseVoiceAgentWebsocket]] - Websocket base class used by voice agents | ||
* [[Run_llama_Llama_index_BaseVoiceAgentInterface]] - Audio I/O interface used by voice agents | * [[Implementation:Run_llama_Llama_index_BaseVoiceAgentInterface]] - Audio I/O interface used by voice agents | ||
* [[Run_llama_Llama_index_Tool_Calling]] - Tool invocation utilities for agent tool use | * [[Implementation:Run_llama_Llama_index_Tool_Calling]] - Tool invocation utilities for agent tool use | ||
[[Category:Implementation]] | [[Category:Implementation]] | ||
Latest revision as of 10:51, 27 September 2026
Overview
The BaseVoiceAgent class is an abstract base class that defines the interface and shared functionality for all voice agents in LlamaIndex. It provides the foundational structure for voice-based conversational AI, combining a websocket connection for voice services, an audio input/output interface, optional tool usage, and conversation state management (messages and events).
Source File: llama-index-core/llama_index/core/voice_agents/base.py
Module: llama_index.core.voice_agents.base
Lines of Code: 162
Dependencies
| Dependency | Type | Purpose |
|---|---|---|
abc.ABC |
Standard Library | Abstract base class support |
abc.abstractmethod |
Standard Library | Decorator for abstract methods |
typing |
Standard Library | Type annotations (Optional, Any, Callable, List)
|
llama_index.core.voice_agents.websocket.BaseVoiceAgentWebsocket |
Internal | Websocket base class for voice services |
llama_index.core.voice_agents.interface.BaseVoiceAgentInterface |
Internal | Audio I/O interface base class |
llama_index.core.voice_agents.events.BaseVoiceAgentEvent |
Internal | Event data model for voice agent events |
llama_index.core.llms.ChatMessage |
Internal | Chat message data model |
llama_index.core.tools.BaseTool |
Internal | Base tool interface for agent tool use |
Class: BaseVoiceAgent
class BaseVoiceAgent(ABC)
Constructor
def __init__(
self,
ws: Optional[BaseVoiceAgentWebsocket] = None,
interface: Optional[BaseVoiceAgentInterface] = None,
ws_url: Optional[str] = None,
api_key: Optional[str] = None,
tools: Optional[List[BaseTool]] = None,
)
| Parameter | Type | Default | Description |
|---|---|---|---|
ws |
Optional[BaseVoiceAgentWebsocket] |
None |
The websocket instance providing the voice service connection |
interface |
Optional[BaseVoiceAgentInterface] |
None |
The audio input/output interface (microphone/speaker) |
ws_url |
Optional[str] |
None |
URL for websocket connection (alternative to passing a ws instance)
|
api_key |
Optional[str] |
None |
API key for authentication with the voice service |
tools |
Optional[List[BaseTool]] |
None |
List of tools the agent can use during conversation |
Instance Attributes
| Attribute | Type | Description |
|---|---|---|
ws |
Optional[BaseVoiceAgentWebsocket] |
Websocket connection for voice service |
ws_url |
Optional[str] |
Websocket URL |
interface |
Optional[BaseVoiceAgentInterface] |
Audio I/O interface |
api_key |
Optional[str] |
API key |
tools |
Optional[List[BaseTool]] |
Available tools |
_messages |
List[ChatMessage] |
Private list of conversation messages (initialized empty) |
_events |
List[BaseVoiceAgentEvent] |
Private list of conversation events (initialized empty) |
Abstract Methods
All abstract methods are asynchronous, reflecting the real-time nature of voice interactions.
start
@abstractmethod async def start(self, *args: Any, **kwargs: Any) -> None
Starts the voice agent. Subclasses must implement initialization logic such as establishing websocket connections and starting audio streams.
send
@abstractmethod async def send(self, audio: Any, *args: Any, **kwargs: Any) -> None
Sends audio data to the websocket underlying the voice agent. The audio parameter is typed as Any to accommodate various audio formats (bytes, strings, etc.).
handle_message
@abstractmethod async def handle_message(self, message: Any, *args: Any, **kwargs: Any) -> Any
Handles an incoming message from the voice service. The message is typically a dictionary but is typed as Any for flexibility. Returns any output that may result from processing the message.
interrupt
@abstractmethod async def interrupt(self) -> None
Interrupts the current audio input/output flow. Used for scenarios like barge-in (when the user starts speaking while the agent is still outputting audio).
stop
@abstractmethod async def stop(self) -> None
Stops the voice agent conversation. Subclasses should implement cleanup logic such as closing websocket connections and stopping audio streams.
Concrete Methods
export_messages
def export_messages(
self,
limit: Optional[int] = None,
filter: Optional[Callable[[List[ChatMessage]], List[ChatMessage]]] = None,
) -> List[ChatMessage]
Exports recorded conversation messages with optional limiting and filtering.
Logic:
- Starts with the full
_messageslist. - If
limitis provided and is less than or equal to the number of messages, truncates to the firstlimitmessages. - If a
filtercallable is provided, applies it to the message list. - Returns the resulting list.
| Parameter | Type | Description |
|---|---|---|
limit |
Optional[int] |
Maximum number of messages to return |
filter |
Optional[Callable] |
A function that takes a list of messages and returns a filtered list |
export_events
def export_events(
self,
limit: Optional[int] = None,
filter: Optional[Callable[[List[BaseVoiceAgentEvent]], List[BaseVoiceAgentEvent]]] = None,
) -> List[BaseVoiceAgentEvent]
Exports recorded conversation events with optional limiting and filtering. Follows the same logic as export_messages but operates on _events instead.
| Parameter | Type | Description |
|---|---|---|
limit |
Optional[int] |
Maximum number of events to return |
filter |
Optional[Callable] |
A function that takes a list of events and returns a filtered list |
Architecture
BaseVoiceAgent (abstract)
|
+--- ws: BaseVoiceAgentWebsocket -- handles websocket communication
|
+--- interface: BaseVoiceAgentInterface -- handles audio I/O (mic/speaker)
|
+--- tools: List[BaseTool] -- optional tool use
|
+--- _messages: List[ChatMessage] -- conversation history
|
+--- _events: List[BaseVoiceAgentEvent] -- event log
|
+--- Abstract methods: start(), send(), handle_message(), interrupt(), stop()
|
+--- Concrete methods: export_messages(), export_events()
Design Patterns
Template Method Pattern
The class defines the skeleton of a voice agent's lifecycle (start, send, handle_message, interrupt, stop) as abstract methods, while providing concrete utility methods (export_messages, export_events) for state inspection. Subclasses fill in the implementation details specific to their voice service provider.
Composition Over Inheritance
The agent composes a websocket (BaseVoiceAgentWebsocket) and an audio interface (BaseVoiceAgentInterface) rather than inheriting from them. This allows flexible mixing of different websocket implementations and audio interfaces.
Event Sourcing
The _messages and _events lists maintain a complete history of the conversation. The export methods support filtering and limiting for downstream analysis or debugging.
See Also
- Implementation:Run_llama_Llama_index_BaseVoiceAgentWebsocket - Websocket base class used by voice agents
- Implementation:Run_llama_Llama_index_BaseVoiceAgentInterface - Audio I/O interface used by voice agents
- Implementation:Run_llama_Llama_index_Tool_Calling - Tool invocation utilities for agent tool use