client
The harness’s connection to the bot under test.
EvalClient runs the eval pipeline, a Pipecat pipeline acting as an
RTVI client: it waits for the bot to listen, connects, completes the
handshake, sends the user’s turns (text, synthesized speech, a recording,
DTMF keys, an image), records the conversation, and tears down. The bot’s
output reaches the rest of the harness as events on the stream the client
is given. In a simulation the pipeline also carries the persona LLM, which
answers the bot on its own.
- class pipecat.evals.client.EvalClientParams(*, bot_audio: bool = False, user_audio: bool = False, user_speech: dict | None = None, capture_bot_audio: bool = False, capture_bot_images: bool = False, report_level: str | None = None, vad_events: bool = False, marker_events: bool = False, context: list[dict] = <factory>, trigger_disconnect: bool = False)[source]
Bases:
BaseModelWhat the scenario asks of the bot; the session derives it from its scenario.
- Parameters:
bot_audio – Whether the bot speaks; in text mode it is asked to skip TTS at connect.
user_audio – Whether the user’s turns reach the bot as audio.
user_speech – TTS config the user’s turns are synthesized with, or
None; sets the user audio rate.capture_bot_audio – Whether the bot forwards its synthesized audio, for the
responsetranscription or a persona that listens.capture_bot_images – Whether the bot reports the images it outputs.
report_level – Function-call report level to ask of the bot, or
Nonefor its default.vad_events – Whether to ask the bot for its raw VAD events.
marker_events – Whether to ask the bot for the sideband markers its LLM emits.
context – Messages the bot’s context starts from, sent right after the handshake; empty sends nothing.
trigger_disconnect – Whether the scenario itself asks for the bot’s
on_client_disconnectedhandler to fire when this connection ends; the run’sEvalSessionParamscan ask too.
- class pipecat.evals.client.EvalClient(bot_url: str, *, params: EvalClientParams | None = None, session_params: EvalSessionParams | None = None, stream: EvalEventStream, trace: EvalTrace, user_tts: CachingTTSService | None = None, bot_stt: STTService | None = None, persona: EvalPersona | None = None)[source]
Bases:
objectThe eval pipeline that talks to the bot, and the user’s sends into it.
Built once per run:
wait_for_bot(),start(),handshake(), the turns (send_text(),say(),play(),send_dtmf(),send_image()), thenstop(). The pipeline is the same for both modalities and both kinds of eval:input -> [STT -> speech gate -> user aggregator] -> sink -> [persona LLM] -> [user TTS] -> [persona relay] -> output -> [assistant aggregator]
The bracketed stages exist only in audio mode (the STT, the gate that drops transcripts of audio from before a send that talked over the bot, the aggregator, the user TTS) or in a simulation (the persona LLM, its relay, the aggregator that records its replies). The bot’s output comes in through the input and stops at the sink, as events; the user’s turns start at the sink and leave through the output. A session builds it, with the params its scenario asks for.
- __init__(bot_url: str, *, params: EvalClientParams | None = None, session_params: EvalSessionParams | None = None, stream: EvalEventStream, trace: EvalTrace, user_tts: CachingTTSService | None = None, bot_stt: STTService | None = None, persona: EvalPersona | None = None)[source]
Initialize the client.
- Parameters:
bot_url – WebSocket URL of the bot’s eval transport.
params – What the scenario asks of the bot;
Noneasks nothing beyond a text-mode conversation.session_params – How the run behaves: the connect timeout, the recording, and the teardown are the client’s part of it.
Nonefor the defaults.stream – Where the bot’s output is appended as events.
trace – The run’s trace.
user_tts – The
CachingTTSServicethat synthesizes spoken user turns, orNonefor text mode.bot_stt – The
STTServicethat transcribes the bot’s audio into theresponseevent, orNonewhen unused.persona – A simulation’s
EvalPersona, whose LLM rides in the pipeline and whose context the aggregators keep up to date with both sides of the conversation;Nonefor a scripted scenario.
- property has_user_tts: bool
Whether text user turns are spoken to the bot rather than sent as text.
- property sends_user_audio: bool
Whether the user’s side streams audio to the bot (the output is enabled).
- async wait_for_bot() None[source]
Wait until the bot accepts connections.
Retries a plain TCP connect until the bot is listening, and stops short of the WebSocket handshake on purpose: a full connection would make the bot greet, and that connection would then be thrown away.
- Raises:
OSError – The last connect error, when the bot never accepted within the connect timeout.
TimeoutError – When the timeout passed without a single attempt failing.
- async start() None[source]
Build the eval pipeline and start running it, which connects to the bot.
- async handshake() None[source]
Wait for
bot-ready, then send what the scenario asks of the bot and its context.A bot that never says ready is not a usable eval target, so this raises instead of sending turns to a half-started bot.
- Raises:
TimeoutError – If the bot never sends
bot-ready.
- async send(message: Message) None[source]
Send an RTVI client message to the bot through the transport pipeline.
- Parameters:
message – The message to send.
- async send_text(text: str) None[source]
Send a text user turn via the RTVI
send-textmessage.- Parameters:
text – The user’s turn.
- async send_dtmf(keys: str) None[source]
Send a DTMF keypress turn as one RTVI
dtmfmessage.The bot handles the keys the way it handles a phone keypad.
- Parameters:
keys – The keys to press, in order.
- async send_image(image_path: str) None[source]
Register an image for the current turn.
The bot’s eval transport serves it back when the bot asks for a user image. The file is sent as is.
- Parameters:
image_path – Path to the image file.
- async say(text: str) None[source]
Speak
textas the user.The user TTS synthesizes it (cached) and the output paces it to the bot as live audio.
- Parameters:
text – What the user says.
- async play(path: str) None[source]
Play a recording to the bot as the user’s turn, instead of synthesizing one.
It goes out like a spoken turn: resampled to the user audio rate, paced to the bot, and recorded.
- Parameters:
path – Path to the audio file.