session
The eval session: one conversation with a bot, driven to a result.
A session connects to a running bot’s eval transport as an RTVI client,
runs the handshake, lets its driver converse, tears down, and returns the
driver’s result, a failed connect or a harness error included. The two
session kinds build the client and the driver for their kind of scenario;
EvalSession.from_scenario() builds whichever kind a scenario is, and
EvalSessionParams is how the run behaves, whichever kind it is.
Example:
params = EvalSessionParams(stop_bot=True)
for scenario in EvalScenarioFile.load("scenarios/greeting.yaml"):
session = EvalSession.from_scenario(scenario, "ws://localhost:7860", params=params)
result = await session.run()
print(scenario.name, "PASS" if result.passed else "FAIL")
- class pipecat.evals.session.EvalSessionParams(*, connect_timeout_s: float = 5.0, default_timeout_ms: int = 60000, record_path: str | None = None, cache_dir: str | None = None, use_cache: bool = True, stop_bot: bool = False, trigger_disconnect: bool = False)[source]
Bases:
BaseModelHow a run behaves, whichever kind of scenario it is: timeouts, recording, caching, teardown.
Plain configuration, so one instance serves many runs and crosses process boundaries; the services a run uses are passed to the session separately.
- Parameters:
connect_timeout_s – How long to wait for the bot to accept the WS connection before giving up.
default_timeout_ms – Scripted scenarios only: the latency budget for expectations without their own
within_ms(the turn’s expectations share one deadline anchored at the send). Defaults to 60s.record_path – Where to save the conversation audio, or
None. Only an audio-mode run records.cache_dir – Directory for cached synthesized user audio, or
Nonefor the default (<user-cache-dir>/pipecat/evals/tts).use_cache – When False, ignore cached user audio and force fresh synthesis, with no cache reads or writes.
stop_bot – When True, ask the bot to cancel its pipeline, and exit, on teardown. Leave False to keep it running for more scenarios.
trigger_disconnect – When True, fire the bot’s
on_client_disconnectedhandler when the connection ends. A scenario’s owntrigger_disconnectfield also opts in. Bots often cancel their pipeline there, so it is off by default to avoid that between scenarios.
- class pipecat.evals.session.EvalSession(*, kind: EvalKind, name: str, bot_url: str, params: EvalSessionParams | None = None)[source]
Bases:
BaseObject,Generic[R]One conversation with a bot, driven to a result.
Connect, run the handshake, let the driver converse, tear down, and turn what happened, a failed connect or a harness error included, into the driver’s result. The subclasses,
EvalScriptSessionandEvalSimulationSession, build the client and the driver for their kind of scenario;from_scenario()picks the one a scenario needs.Event handlers available:
on_progress: Called with an
EvalProgressrecord as the conversation advances: a scripted scenario’s turns and expectations as they resolve, a simulation’s lines as they are spoken. Records are emitted in order, andrun()waits for every handler before it returns.
- __init__(*, kind: EvalKind, name: str, bot_url: str, params: EvalSessionParams | None = None)[source]
Initialize the session’s runtime.
- Parameters:
kind – The scenario kind being run, for the trace.
name – The scenario’s or simulation’s name.
bot_url – WebSocket URL of the bot’s eval transport.
params – How the run behaves;
Nonefor the defaults.
- classmethod from_scenario(scenario: EvalScriptScenario, bot_url: str, *, params: EvalSessionParams | None = None, judge: EvalJudge | None = None, user_tts: CachingTTSService | None = None, bot_stt: STTService | None = None, on_progress: Callable[[EvalScriptTurnProgress], None] | None = None, connect_timeout_s: float | None = None, default_timeout_ms: int | None = None, record_path: str | None = None, cache_dir: str | None = None, use_cache: bool | None = None, stop_bot: bool | None = None, trigger_disconnect: bool | None = None) EvalScriptSession[source]
- classmethod from_scenario(scenario: EvalSimulationScenario, bot_url: str, *, params: EvalSessionParams | None = None, persona_llm: LLMService | None = None, judge: EvalJudge | None = None, user_tts: CachingTTSService | None = None, bot_stt: STTService | None = None, connect_timeout_s: float | None = None, default_timeout_ms: int | None = None, record_path: str | None = None, cache_dir: str | None = None, use_cache: bool | None = None, stop_bot: bool | None = None, trigger_disconnect: bool | None = None) EvalSimulationSession
- classmethod from_scenario(scenario: EvalScriptScenario | EvalSimulationScenario, bot_url: str, *, params: EvalSessionParams | None = None, persona_llm: LLMService | None = None, judge: EvalJudge | None = None, user_tts: CachingTTSService | None = None, bot_stt: STTService | None = None, on_progress: Callable[[EvalScriptTurnProgress], None] | None = None, connect_timeout_s: float | None = None, default_timeout_ms: int | None = None, record_path: str | None = None, cache_dir: str | None = None, use_cache: bool | None = None, stop_bot: bool | None = None, trigger_disconnect: bool | None = None) EvalScriptSession | EvalSimulationSession
Build a ready-to-run session for a scenario of either kind.
A scripted scenario gets an
EvalScriptSession, a simulation anEvalSimulationSession, each constructing the services it needs. Passpersona_llm,judge,user_tts, orbot_sttto use your own. Then awaitrun():session = EvalSession.from_scenario(scenario, "ws://localhost:7860") result = await session.run()
- Parameters:
scenario – The parsed scenario to run, scripted or a simulation.
bot_url – WebSocket URL of the bot’s eval transport.
params – How the run behaves;
Nonefor the defaults.persona_llm – Simulations only: override the persona LLM (default: built from the simulation’s
simulator).judge – Override the judge (default: built from the scenario’s
judgewhen the run needs one).user_tts – Override the user-audio TTS (default: built from the scenario’s
user_speechin audio mode).bot_stt – Override the bot-audio STT (default: built from the scenario’s
transcriberwhen the run transcribes the bot).on_progress –
Scripted scenarios only: a callback for each turn and expectation as it resolves.
Deprecated since version 1.9.0: Use the
on_progressevent handler instead. Will be removed in 2.0.0.connect_timeout_s –
The
paramsfield of the same name.Deprecated since version 1.9.0: Use
paramsinstead. Will be removed in 2.0.0.default_timeout_ms –
The
paramsfield of the same name.Deprecated since version 1.9.0: Use
paramsinstead. Will be removed in 2.0.0.record_path –
The
paramsfield of the same name.Deprecated since version 1.9.0: Use
paramsinstead. Will be removed in 2.0.0.cache_dir –
The
paramsfield of the same name.Deprecated since version 1.9.0: Use
paramsinstead. Will be removed in 2.0.0.use_cache –
The
paramsfield of the same name.Deprecated since version 1.9.0: Use
paramsinstead. Will be removed in 2.0.0.stop_bot –
The
paramsfield of the same name.Deprecated since version 1.9.0: Use
paramsinstead. Will be removed in 2.0.0.trigger_disconnect –
The
paramsfield of the same name.Deprecated since version 1.9.0: Use
paramsinstead. Will be removed in 2.0.0.
- Returns:
A configured session of the scenario’s kind, ready for
run().- Raises:
ValueError – If
persona_llmis given for a scripted scenario, which has no persona, oron_progressfor a simulation, which reports its progress through the event handler alone.TypeError – If
scenariois neither kind.