base_driver

The base driver: what the user says next, and how the run is scored.

A driver runs the conversation over the session’s client and event stream and assembles the result. The scripted driver plays a file’s turns and matches their expectations; the simulation driver lets the persona LLM hold the conversation and judges the whole of it.

class pipecat.evals.base_driver.BaseEvalDriver(*, client: EvalClient, stream: EvalEventStream, judge: EvalJudge | None, trace: EvalTrace, progress: Callable[[EvalScriptTurnProgress | EvalSimulationProgress], Awaitable[None]])[source]

Bases: ABC, Generic[R]

Base class for the drivers: what the user says next, and how the run is scored.

A driver sends through the client, reads the stream, and decides what counts as success. Subclasses implement run() and result().

__init__(*, client: EvalClient, stream: EvalEventStream, judge: EvalJudge | None, trace: EvalTrace, progress: Callable[[EvalScriptTurnProgress | EvalSimulationProgress], Awaitable[None]])[source]

Initialize the driver.

Parameters:
  • client – The connection to the bot, for the user’s sends.

  • stream – The bot’s output as events.

  • judge – The judge for eval: assertions, or None; the user’s turns are added to its conversation so replies are judged in context.

  • trace – The run’s trace.

  • progress – Awaited with an EvalProgress record as the conversation advances.

abstractmethod async run() → list[EvalAssertionFailure][source]

Drive the conversation to its end and return the failures.

async close() → None[source]

Close the judge, once the run has ended.

abstractmethod result(*, failures: list[EvalAssertionFailure], duration_ms: int, events_seen: list[dict], debug_log: list[str], skipped: str | None = None) → R[source]

Assemble the run’s result from what the driver scored and the session saw.

Parameters:
  • failures – The run’s failures: the driver’s own, plus the session’s (a failed connect or handshake, a harness error).

  • duration_ms – Wall-clock time the run took.

  • events_seen – Every event observed, for diagnostics.

  • debug_log – The run’s trace.

  • skipped – Why the run was not driven at all, or None.

record_failure(failure: EvalAssertionFailure) → None[source]

Note a run-level failure the session raised while the driver was running.

Parameters:

failure – The failure, scored against the trace’s current turn.