session

The eval session: one conversation with a bot, driven to a result.

A session connects to a running bot’s eval transport as an RTVI client, runs the handshake, lets its driver converse, tears down, and returns the driver’s result, a failed connect or a harness error included. The two session kinds build the client and the driver for their kind of scenario; EvalSession.from_scenario() builds whichever kind a scenario is, and EvalSessionParams is how the run behaves, whichever kind it is.

Example:

params = EvalSessionParams(stop_bot=True)
for scenario in EvalScenarioFile.load("scenarios/greeting.yaml"):
    session = EvalSession.from_scenario(scenario, "ws://localhost:7860", params=params)
    result = await session.run()
    print(scenario.name, "PASS" if result.passed else "FAIL")
class pipecat.evals.session.EvalSessionParams(*, connect_timeout_s: float = 5.0, default_timeout_ms: int = 60000, record_path: str | None = None, cache_dir: str | None = None, use_cache: bool = True, stop_bot: bool = False, trigger_disconnect: bool = False)[source]

Bases: BaseModel

How a run behaves, whichever kind of scenario it is: timeouts, recording, caching, teardown.

Plain configuration, so one instance serves many runs and crosses process boundaries; the services a run uses are passed to the session separately.

Parameters:
  • connect_timeout_s – How long to wait for the bot to accept the WS connection before giving up.

  • default_timeout_ms – Scripted scenarios only: the latency budget for expectations without their own within_ms (the turn’s expectations share one deadline anchored at the send). Defaults to 60s.

  • record_path – Where to save the conversation audio, or None. Only an audio-mode run records.

  • cache_dir – Directory for cached synthesized user audio, or None for the default (<user-cache-dir>/pipecat/evals/tts).

  • use_cache – When False, ignore cached user audio and force fresh synthesis, with no cache reads or writes.

  • stop_bot – When True, ask the bot to cancel its pipeline, and exit, on teardown. Leave False to keep it running for more scenarios.

  • trigger_disconnect – When True, fire the bot’s on_client_disconnected handler when the connection ends. A scenario’s own trigger_disconnect field also opts in. Bots often cancel their pipeline there, so it is off by default to avoid that between scenarios.

class pipecat.evals.session.EvalSession(*, kind: EvalKind, name: str, bot_url: str, params: EvalSessionParams | None = None)[source]

Bases: BaseObject, Generic[R]

One conversation with a bot, driven to a result.

Connect, run the handshake, let the driver converse, tear down, and turn what happened, a failed connect or a harness error included, into the driver’s result. The subclasses, EvalScriptSession and EvalSimulationSession, build the client and the driver for their kind of scenario; from_scenario() picks the one a scenario needs.

Event handlers available:

  • on_progress: Called with an EvalProgress record as the conversation advances: a scripted scenario’s turns and expectations as they resolve, a simulation’s lines as they are spoken. Records are emitted in order, and run() waits for every handler before it returns.

__init__(*, kind: EvalKind, name: str, bot_url: str, params: EvalSessionParams | None = None)[source]

Initialize the session’s runtime.

Parameters:
  • kind – The scenario kind being run, for the trace.

  • name – The scenario’s or simulation’s name.

  • bot_url – WebSocket URL of the bot’s eval transport.

  • params – How the run behaves; None for the defaults.

classmethod from_scenario(scenario: EvalScriptScenario, bot_url: str, *, params: EvalSessionParams | None = None, judge: EvalJudge | None = None, user_tts: CachingTTSService | None = None, bot_stt: STTService | None = None, on_progress: Callable[[EvalScriptTurnProgress], None] | None = None, connect_timeout_s: float | None = None, default_timeout_ms: int | None = None, record_path: str | None = None, cache_dir: str | None = None, use_cache: bool | None = None, stop_bot: bool | None = None, trigger_disconnect: bool | None = None) → EvalScriptSession[source]
classmethod from_scenario(scenario: EvalSimulationScenario, bot_url: str, *, params: EvalSessionParams | None = None, persona_llm: LLMService | None = None, judge: EvalJudge | None = None, user_tts: CachingTTSService | None = None, bot_stt: STTService | None = None, connect_timeout_s: float | None = None, default_timeout_ms: int | None = None, record_path: str | None = None, cache_dir: str | None = None, use_cache: bool | None = None, stop_bot: bool | None = None, trigger_disconnect: bool | None = None) → EvalSimulationSession
classmethod from_scenario(scenario: EvalScriptScenario | EvalSimulationScenario, bot_url: str, *, params: EvalSessionParams | None = None, persona_llm: LLMService | None = None, judge: EvalJudge | None = None, user_tts: CachingTTSService | None = None, bot_stt: STTService | None = None, on_progress: Callable[[EvalScriptTurnProgress], None] | None = None, connect_timeout_s: float | None = None, default_timeout_ms: int | None = None, record_path: str | None = None, cache_dir: str | None = None, use_cache: bool | None = None, stop_bot: bool | None = None, trigger_disconnect: bool | None = None) → EvalScriptSession | EvalSimulationSession

Build a ready-to-run session for a scenario of either kind.

A scripted scenario gets an EvalScriptSession, a simulation an EvalSimulationSession, each constructing the services it needs. Pass persona_llm, judge, user_tts, or bot_stt to use your own. Then await run():

session = EvalSession.from_scenario(scenario, "ws://localhost:7860")
result = await session.run()
Parameters:
  • scenario – The parsed scenario to run, scripted or a simulation.

  • bot_url – WebSocket URL of the bot’s eval transport.

  • params – How the run behaves; None for the defaults.

  • persona_llm – Simulations only: override the persona LLM (default: built from the simulation’s simulator).

  • judge – Override the judge (default: built from the scenario’s judge when the run needs one).

  • user_tts – Override the user-audio TTS (default: built from the scenario’s user_speech in audio mode).

  • bot_stt – Override the bot-audio STT (default: built from the scenario’s transcriber when the run transcribes the bot).

  • on_progress –

    Scripted scenarios only: a callback for each turn and expectation as it resolves.

    Deprecated since version 1.9.0: Use the on_progress event handler instead. Will be removed in 2.0.0.

  • connect_timeout_s –

    The params field of the same name.

    Deprecated since version 1.9.0: Use params instead. Will be removed in 2.0.0.

  • default_timeout_ms –

    The params field of the same name.

    Deprecated since version 1.9.0: Use params instead. Will be removed in 2.0.0.

  • record_path –

    The params field of the same name.

    Deprecated since version 1.9.0: Use params instead. Will be removed in 2.0.0.

  • cache_dir –

    The params field of the same name.

    Deprecated since version 1.9.0: Use params instead. Will be removed in 2.0.0.

  • use_cache –

    The params field of the same name.

    Deprecated since version 1.9.0: Use params instead. Will be removed in 2.0.0.

  • stop_bot –

    The params field of the same name.

    Deprecated since version 1.9.0: Use params instead. Will be removed in 2.0.0.

  • trigger_disconnect –

    The params field of the same name.

    Deprecated since version 1.9.0: Use params instead. Will be removed in 2.0.0.

Returns:

A configured session of the scenario’s kind, ready for run().

Raises:
  • ValueError – If persona_llm is given for a scripted scenario, which has no persona, or on_progress for a simulation, which reports its progress through the event handler alone.

  • TypeError – If scenario is neither kind.

async run() → R[source]

Connect, drive the conversation, and return the result.