simulation_driver

The simulation driver: lets the persona hold the conversation, then judges the whole of it.

class pipecat.evals.simulation_driver.EvalSimulationDriver(*, simulation: EvalSimulationScenario, persona: EvalPersona, client: EvalClient, stream: EvalEventStream, judge: EvalJudge | None, trace: EvalTrace, progress: Callable[[EvalScriptTurnProgress | EvalSimulationProgress], Awaitable[None]])[source]

Bases: BaseEvalDriver[EvalSimulationResult]

Lets the persona LLM hold the conversation, then judges the whole of it.

The persona answers the bot on its own inside the client’s pipeline. The driver reports each line as progress and watches for the end: the persona’s end_call, the bot hanging up, the turn cap, the time cap, a lull neither side breaks, or a failure of the harness’s own pipeline. Then one judge call over the whole transcript scores every bot turn on every criterion and decides the goal.

__init__(*, simulation: EvalSimulationScenario, persona: EvalPersona, client: EvalClient, stream: EvalEventStream, judge: EvalJudge | None, trace: EvalTrace, progress: Callable[[EvalScriptTurnProgress | EvalSimulationProgress], Awaitable[None]])[source]

Initialize the driver.

Parameters:
  • simulation – The simulation being run.

  • persona – The simulated caller: its instruction, its LLM in the pipeline, and its context.

  • client – The connection to the bot.

  • stream – The bot’s output as events.

  • judge – The judge for the goal and the quality criteria.

  • trace – The run’s trace.

  • progress – Awaited with an EvalSimulationProgress for each line of the conversation, and once when it ends.

async run() → list[EvalAssertionFailure][source]

Watch the conversation until it ends, then judge it.

result(*, failures: list[EvalAssertionFailure], duration_ms: int, events_seen: list[dict], debug_log: list[str], skipped: str | None = None) → EvalSimulationResult[source]

The run’s result; a run-level failure makes it an error, not a goal failure.

timeline() → list[dict][source]

The conversation as the events told it, each line with the tool calls made by then.

A bot turn is everything the bot said between two persona turns, however turn detection or a function call split it; a turn in which the bot said nothing is not a turn. Built as the events arrive.

transcript() → list[dict][source]

The conversation as the judge sees it: the lines, each tool call in place before the line it preceded.

Deprecated since version 1.11.0: No replacement: the judge keeps the conversation it judges. Will be removed in 2.0.0.

conversation() → list[dict][source]

The conversation with the persona as user and the bot as assistant, without the tool calls.

tool_calls() → list[str][source]

The bot’s function calls in order, one line each; a cancelled call is listed as cancelled.