script_driver

The scripted driver: plays a scenario’s turns and matches each turn’s expectations.

class pipecat.evals.script_driver.EvalScriptDriver(*, scenario: EvalScriptScenario, default_timeout_ms: int, client: EvalClient, stream: EvalEventStream, judge: EvalJudge | None, trace: EvalTrace, progress: Callable[[EvalScriptTurnProgress | EvalSimulationProgress], Awaitable[None]])[source]

Bases: BaseEvalDriver[EvalScriptResult]

Plays a scenario’s turns in order and matches each turn’s expectations.

A failed turn ends the scenario by default, since the conversation is in an unknown state from there on; stop_on_failure: false drives every turn regardless.

__init__(*, scenario: EvalScriptScenario, default_timeout_ms: int, client: EvalClient, stream: EvalEventStream, judge: EvalJudge | None, trace: EvalTrace, progress: Callable[[EvalScriptTurnProgress | EvalSimulationProgress], Awaitable[None]])[source]

Initialize the driver.

Parameters:
  • scenario – The scenario whose turns to play.

  • default_timeout_ms – Latency budget for expectations without their own within_ms.

  • client – The connection to the bot.

  • stream – The bot’s output as events.

  • judge – The judge for eval: assertions, or None.

  • trace – The run’s trace.

  • progress – Awaited with an EvalProgress as turns and expectations resolve.

async run() → list[EvalAssertionFailure][source]

Drive the scenario’s turns in order, filling in their records.

result(*, failures: list[EvalAssertionFailure], duration_ms: int, events_seen: list[dict], debug_log: list[str], skipped: str | None = None) → EvalScriptResult[source]

The scenario’s result: passed only if nothing failed and nothing was skipped.

record_failure(failure: EvalAssertionFailure) → None[source]

Score a run-level failure against the turn it interrupted, if any.

Before any turn started (a sub-pipeline that never came up) the failure’s turn is -1 and every turn stays not_run.