client

The harness’s connection to the bot under test.

EvalClient runs the eval pipeline, a Pipecat pipeline acting as an RTVI client: it waits for the bot to listen, connects, completes the handshake, sends the user’s turns (text, synthesized speech, a recording, DTMF keys, an image), records the conversation, and tears down. The bot’s output reaches the rest of the harness as events on the stream the client is given. In a simulation the pipeline also carries the persona LLM, which answers the bot on its own.

class pipecat.evals.client.EvalClientParams(*, bot_audio: bool = False, user_audio: bool = False, user_speech: dict | None = None, capture_bot_audio: bool = False, capture_bot_images: bool = False, report_level: str | None = None, vad_events: bool = False, marker_events: bool = False, context: list[dict] = <factory>, trigger_disconnect: bool = False)[source]

Bases: BaseModel

What the scenario asks of the bot; the session derives it from its scenario.

Parameters:
  • bot_audio – Whether the bot speaks; in text mode it is asked to skip TTS at connect.

  • user_audio – Whether the user’s turns reach the bot as audio.

  • user_speech – TTS config the user’s turns are synthesized with, or None; sets the user audio rate.

  • capture_bot_audio – Whether the bot forwards its synthesized audio, for the response transcription or a persona that listens.

  • capture_bot_images – Whether the bot reports the images it outputs.

  • report_level – Function-call report level to ask of the bot, or None for its default.

  • vad_events – Whether to ask the bot for its raw VAD events.

  • marker_events – Whether to ask the bot for the sideband markers its LLM emits.

  • context – Messages the bot’s context starts from, sent right after the handshake; empty sends nothing.

  • trigger_disconnect – Whether the scenario itself asks for the bot’s on_client_disconnected handler to fire when this connection ends; the run’s EvalSessionParams can ask too.

class pipecat.evals.client.EvalClient(bot_url: str, *, params: EvalClientParams | None = None, session_params: EvalSessionParams | None = None, stream: EvalEventStream, trace: EvalTrace, user_tts: CachingTTSService | None = None, bot_stt: STTService | None = None, persona: EvalPersona | None = None)[source]

Bases: object

The eval pipeline that talks to the bot, and the user’s sends into it.

Built once per run: wait_for_bot(), start(), handshake(), the turns (send_text(), say(), play(), send_dtmf(), send_image()), then stop(). The pipeline is the same for both modalities and both kinds of eval:

input -> [STT -> speech gate -> user aggregator] -> sink -> [persona LLM]
      -> [user TTS] -> [persona relay] -> output -> [assistant aggregator]

The bracketed stages exist only in audio mode (the STT, the gate that drops transcripts of audio from before a send that talked over the bot, the aggregator, the user TTS) or in a simulation (the persona LLM, its relay, the aggregator that records its replies). The bot’s output comes in through the input and stops at the sink, as events; the user’s turns start at the sink and leave through the output. A session builds it, with the params its scenario asks for.

__init__(bot_url: str, *, params: EvalClientParams | None = None, session_params: EvalSessionParams | None = None, stream: EvalEventStream, trace: EvalTrace, user_tts: CachingTTSService | None = None, bot_stt: STTService | None = None, persona: EvalPersona | None = None)[source]

Initialize the client.

Parameters:
  • bot_url – WebSocket URL of the bot’s eval transport.

  • params – What the scenario asks of the bot; None asks nothing beyond a text-mode conversation.

  • session_params – How the run behaves: the connect timeout, the recording, and the teardown are the client’s part of it. None for the defaults.

  • stream – Where the bot’s output is appended as events.

  • trace – The run’s trace.

  • user_tts – The CachingTTSService that synthesizes spoken user turns, or None for text mode.

  • bot_stt – The STTService that transcribes the bot’s audio into the response event, or None when unused.

  • persona – A simulation’s EvalPersona, whose LLM rides in the pipeline and whose context the aggregators keep up to date with both sides of the conversation; None for a scripted scenario.

property has_user_tts: bool

Whether text user turns are spoken to the bot rather than sent as text.

property sends_user_audio: bool

Whether the user’s side streams audio to the bot (the output is enabled).

async wait_for_bot() → None[source]

Wait until the bot accepts connections.

Retries a plain TCP connect until the bot is listening, and stops short of the WebSocket handshake on purpose: a full connection would make the bot greet, and that connection would then be thrown away.

Raises:
  • OSError – The last connect error, when the bot never accepted within the connect timeout.

  • TimeoutError – When the timeout passed without a single attempt failing.

async start() → None[source]

Build the eval pipeline and start running it, which connects to the bot.

async handshake() → None[source]

Wait for bot-ready, then send what the scenario asks of the bot and its context.

A bot that never says ready is not a usable eval target, so this raises instead of sending turns to a half-started bot.

Raises:

TimeoutError – If the bot never sends bot-ready.

async stop() → None[source]

Save the recording, optionally cancel the bot, and end the pipeline.

async send(message: Message) → None[source]

Send an RTVI client message to the bot through the transport pipeline.

Parameters:

message – The message to send.

async send_text(text: str) → None[source]

Send a text user turn via the RTVI send-text message.

Parameters:

text – The user’s turn.

async send_dtmf(keys: str) → None[source]

Send a DTMF keypress turn as one RTVI dtmf message.

The bot handles the keys the way it handles a phone keypad.

Parameters:

keys – The keys to press, in order.

async send_image(image_path: str) → None[source]

Register an image for the current turn.

The bot’s eval transport serves it back when the bot asks for a user image. The file is sent as is.

Parameters:

image_path – Path to the image file.

async say(text: str) → None[source]

Speak text as the user.

The user TTS synthesizes it (cached) and the output paces it to the bot as live audio.

Parameters:

text – What the user says.

async play(path: str) → None[source]

Play a recording to the bot as the user’s turn, instead of synthesizing one.

It goes out like a spoken turn: resampled to the user audio rate, paced to the bot, and recorded.

Parameters:

path – Path to the audio file.

async configure_persona(instruction: str) → None[source]

Give the persona LLM its instruction, and say whether its replies are spoken.

In text mode a reply goes to the bot as text; in audio mode the user TTS speaks it.

Parameters:

instruction – The persona’s system instruction.

async hang_up() → None[source]

End the persona’s part of the conversation: it answers nothing more.