suite

Multi-bot eval suite runner.

An EvalManifest lists bots to spawn and the scenarios to run against each, scripted and simulated alike. An EvalSuite spawns each bot with its eval transport on its own port and drives it with the harness in a subprocess, several at a time. pipecat eval suite is the CLI in front of it; the release evals are a manifest plus that command.

Manifest format (YAML):

concurrency: 4
repeat: 1                     # run each (bot, scenario) N times
runs_dir: test-runs           # logs + recordings go to <runs_dir>/<timestamp>/
record: false                 # record conversation audio
cache_dir: null               # optional
scenarios_dir: scenarios      # resolved relative to this manifest file
# {python}=interpreter (default sys.executable), {bot}=bot path,
# {port}=assigned per run by the suite runner
spawn: "{python} {bot} -t eval --port {port}"
suite:
  - bot: examples/voice/voice-cartesia.py
    scenarios: [simple_math, greeting]
  - bot: examples/voice/voice-openai.py
    scenarios: [simple_math, interruption]
  - bot: examples/voice/voice-groq.py
    concurrency: 2                           # this bot's own cap
    scenarios: [simple_math, interruption]
  - bot: examples/vision/vision-openai.py
    runner_body:
      path: scenarios/vision-cat.yaml        # passed to the bot as --runner-body
    scenarios: [vision_describe]
  - bot: examples/turns/filter-incomplete-turns.py
    name: openai/gpt-4o-mini                 # this entry's label, the bot path by default
    runner_body:
      data: {model: gpt-4o-mini}             # written to a file for the bot
    scenarios: [turn_completion]
  - bot: examples/turns/filter-incomplete-turns.py
    name: groq/llama-3.3-70b
    runner_body:
      data: {model: llama-3.3-70b}
    scenarios: [turn_completion]
  - bot: examples/flows/restaurant_reservation.py
    scenarios: [book_table]                  # a simulation: its file has a persona

A scenarios: entry names a scenario file, and the file contributes one run per scenario it holds, scripted or a simulation as each says (see EvalScenarioFile), named <file name>/<scenario name>. A name resolves under scenarios_dir with .yaml added and may carry a folder, as scripted/greeting; a name ending in .yaml is a path relative to the manifest instead. A simulation runs as many times as its runs: says, and every run must pass.

An optional runner_body: supplies runner-args data the bot would normally receive in a /start request body (e.g. a vision bot’s image path), passed to it as --runner-body. It holds either path:, a YAML or JSON file resolved relative to the manifest, or data:, the body itself as a mapping, which the suite writes to a file among the run’s logs. A bot given a file runs with the file’s directory as its working directory, so relative paths inside the body (like an image) resolve next to the file; a body that holds such paths belongs in a file for that reason.

Deprecated since version 1.11.0: Use runner_body: {path: <file>} instead of a bare runner_body: <file>. Will be removed in 2.0.0.

An entry’s name: is its label: what the display, the -p filter, the results records and the artifact file names use, and what an entry’s own concurrency: is keyed on. It defaults to the bot: path, so it is only needed when several entries share a bot, as when sweeping a model with runner_body:; two entries may not run the same scenario under the same label.

concurrency is how many runs execute at once. The suite keeps that many going, taking the next run from the first entry in manifest order that still has one, so an entry’s scenarios finish together and no slot waits while any entry still has runs. An entry whose provider rate-limits sets its own concurrency:, and never has more than that many runs in flight.

Manifest-relative paths (bot/bots_dir, scenarios_dir, runs_dir) resolve relative to the manifest file, so a manifest is portable; the same values passed as CLI overrides resolve against the working directory.

repeat (or --repeat) runs every (bot, scenario) pair N times, which is how a flaky behavior gets measured rather than sampled: a bot that passes a scenario half the time looks identical to a reliable one in a single pass. Attempts run attempt-major, every entry’s first attempt before any entry’s second, and carry an EvalRun.attempt number that joins their artifact filenames, so no attempt overwrites another’s logs.

pipecat.evals.suite.capture_pipeline_logs(logs_dir: Path, prefix: str, *, name: str, enabled: bool) → Iterator[None][source]

Capture the harness’s logs for one run into a single <prefix>.debug.log.

The logs are buffered in memory and written on exit, one section per pipeline. The run is tagged with prefix and the sink filters on it, so concurrent runs never mix. Writes nothing unless enabled.

Parameters:
  • logs_dir – Directory the <prefix>.debug.log is written to.

  • prefix – Filename stem; also the eval_run id the sink filters on.

  • name – Human test name shown in each section heading.

  • enabled – When False, do nothing and write no file.

class pipecat.evals.suite.EvalRun(bot: str, scenario: str, scenario_path: Path, name: str | None = None, loaded: EvalScriptScenario | EvalSimulationScenario | None = None, bot_path: Path | None = None, bot_url: str | None = None, runner_body_path: Path | None = None, runner_body: dict | None = None, concurrency: int | None = None, kind: EvalKind = EvalKind.SCRIPT, attempts: int = 1, sweep: bool = False, attempt: int = 1, status: str = 'pending', stopping: bool = False, result: EvalScriptResult | EvalSimulationResult | None = None, error: str | None = None, started_at: float | None = None, duration_ms: int | None = None)[source]

Bases: object

Mutable per-(bot, scenario) state, updated in place so a live display can read it.

Parameters:
  • bot – The manifest’s bot: path (suite) or the bot URL (run).

  • name – The manifest entry’s name:, or None when it has none; label is what the display, the filters and the results use.

  • scenario – Display name (the scenario or simulation, without .yaml).

  • scenario_path – Path to the scenario or simulation file.

  • kind – script (played by EvalScriptSession) or simulation (EvalSimulationSession).

  • attempts – How many times this (bot, scenario) pair runs: the manifest’s repeat, or a simulation’s own runs. Above 1, each attempt’s artifacts carry its number.

  • sweep – Whether the attempts come from a repeat (a measurement: the suite reports a rate and a failure is data) rather than from a simulation’s runs (a requirement: every attempt must pass).

  • bot_path – The bot to spawn (suite); None when connecting to bot_url.

  • bot_url – Connect here instead of spawning (used by pipecat eval run).

  • runner_body_path – Optional --runner-body file for the bot’s runner args.

  • runner_body – The bot’s runner-args body given inline, written to a file for the bot when it is spawned; None when there is none or it comes from runner_body_path.

  • concurrency – How many runs of this bot’s entry may be in flight at once, when its manifest entry says; None is as many as the suite runs.

  • attempt – 1-based attempt number when the suite repeats (see EvalManifest.repeat); always 1 for a single pass.

  • status – pending, running, or done.

  • stopping – True while the bot is being stopped after the run is done. The run still counts as in flight, so the next one waits.

  • result – The outcome, once the run is done.

  • error – Spawn/connection error message, if the run failed before producing a result.

  • started_at – Monotonic start time, for the live elapsed counter.

  • duration_ms – Wall-clock time the run took, in milliseconds.

  • loaded – The scenario, when the run was built from a loaded file; load() returns it, or reads the file when it is None.

name: str | None = None
loaded: EvalScriptScenario | EvalSimulationScenario | None = None
bot_path: Path | None = None
bot_url: str | None = None
runner_body_path: Path | None = None
runner_body: dict | None = None
concurrency: int | None = None
kind: EvalKind = 'script'
attempts: int = 1
sweep: bool = False
attempt: int = 1
status: str = 'pending'
stopping: bool = False
result: EvalScriptResult | EvalSimulationResult | None = None
error: str | None = None
started_at: float | None = None
duration_ms: int | None = None
property label: str

the entry’s name, or its bot when it has none.

Type:

The run’s label

property stem: str

a group entry’s / becomes __.

Type:

The scenario name as a file name stem

load() → EvalScriptScenario | EvalSimulationScenario[source]

The run’s scenario: as loaded when the run was built, else read from its file.

Raises:
class pipecat.evals.suite.EvalManifest(runs: list[EvalRun], spawn: str, python: str, concurrency: int, repeat: int, base_port: int, runs_dir: Path | None, record: bool, cache_dir: str | None)[source]

Bases: object

A parsed eval-suite manifest.

Parameters:
  • runs – The (bot, scenario) and (bot, simulation) runs to execute.

  • spawn – Spawn command template ({python}/{bot}/{port} substituted).

  • python – Interpreter used to spawn each bot.

  • concurrency – How many runs to execute at once.

  • repeat – How many times to run each (bot, scenario) pair. Attempts run attempt-major rather than grouped per bot, so every bot meets the same machine conditions in the same stretch of the sweep and a transient slowdown shows up as a band across all of them instead of a regression in whichever bot happened to be running.

  • base_port – First port to assign; each run gets base_port + index, so the reserved range widens with repeat (bots x scenarios x repeat ports from base_port up).

  • runs_dir – Base for run output (a <name>/ subdir is added), or None.

  • record – Whether to record conversation audio.

  • cache_dir – Directory for cached synthesized user audio, or None.

classmethod load(path: str | Path, *, bots_dir: str | Path | None = None, scenarios_dir: str | Path | None = None, runs_dir: str | Path | None = None, spawn: str | None = None, python: str | None = None, concurrency: int | None = None, repeat: int | None = None, base_port: int | None = None, record: bool | None = None, cache_dir: str | None = None) → EvalManifest[source]

Parse a manifest YAML into an EvalManifest.

A keyword that is not None overrides the manifest’s value, so the CLI wins. Manifest paths resolve against the manifest’s directory, overrides against the working directory.

Parameters:
  • path – Path to the manifest YAML.

  • bots_dir – Override for the manifest’s bots_dir (bot paths are relative to it).

  • scenarios_dir – Override for the manifest’s scenarios_dir.

  • runs_dir – Override for the manifest’s runs_dir (base for run output).

  • spawn – Override for the spawn command template.

  • python – Override for the interpreter used to spawn bots.

  • concurrency – Override for how many runs execute at once.

  • repeat – Override for how many times each (bot, scenario) pair runs. Set here or in the manifest, it also replaces each simulation’s own runs, a repeat of 1 included.

  • base_port – Override for the first port assigned.

  • record – Override for whether to record conversation audio.

  • cache_dir – Override for the synthesized-audio cache directory.

Returns:

The parsed EvalManifest.

class pipecat.evals.suite.EvalSuite(manifest: EvalManifest)[source]

Bases: BaseObject

Runs the (bot, scenario) runs of an EvalManifest, spawning each bot.

Each bot gets its eval transport on its own port and is driven by the harness in a subprocess, several at a time up to the manifest’s concurrency. The runs are updated in place as they go, so a live display can read their progress.

Event handlers available:

  • on_update: Called with an EvalRun whenever that run changes status. Runs are mutated in place, so handlers run synchronously and must return promptly; they see the run as the change left it rather than however it has moved on since.

Example:

manifest = EvalManifest.load("manifest.yaml")
suite = EvalSuite(manifest)
suite.filter(pattern="voice")

@suite.event_handler("on_update")
async def on_update(suite, run):
    print(run.scenario, run.status)

await suite.run(Path("logs"))
__init__(manifest: EvalManifest)[source]

Initialize the suite from a parsed manifest.

Parameters:

manifest – The parsed EvalManifest; its runs become the suite’s working set (narrowed by filter(), executed by run()).

filter(*, pattern: str | None = None, scenario: str | None = None, kind: EvalKind | None = None) → list[EvalRun][source]

Keep only the runs matching a bot substring, a scenario name, and/or a kind.

Parameters:
  • pattern – Keep only runs whose name or bot path contains this substring.

  • scenario – Keep only runs of this scenario: its full <file>/<scenario> name, or either half of it, so a file’s name selects every scenario it holds.

  • kind – Keep only runs of this kind.

Returns:

The matching runs, in their original order.

async run(logs_dir: Path, *, record_dir: Path | None = None, results_path: Path | None = None, on_update: Callable[[EvalRun], None] | None = None, debug: bool = False, params: EvalSessionParams | None = None, use_cache: bool | None = None, default_timeout_ms: int | None = None) → None[source]

Run all of the suite’s runs, in place, with the manifest’s concurrency.

Each run gets its own port (base_port + index). concurrency workers each take the next run and run it until none is left. Runs are taken in manifest order, every entry’s first attempt before any entry’s second, so an entry’s scenarios finish together and no worker waits while any entry still has runs. An entry with a concurrency: of its own never has more than that many runs in flight, across its attempts; a worker that finds it full takes the next entry’s run.

Parameters:
  • logs_dir – Directory for per-run logs.

  • record_dir – Directory for per-run conversation recordings, or None.

  • results_path – JSONL file to append one record per finished run, or None to write none. Each line is flushed as its run completes, so an interrupted sweep keeps everything already finished.

  • on_update –

    Called whenever a run changes status, for live display.

    Deprecated since version 1.9.0: Use the on_update event handler instead. Will be removed in 2.0.0.

  • debug – When True, save each run’s combined <run>.debug.log.

  • params – How each run behaves; None for the defaults. The suite sets what it owns on each run’s copy: the connect timeout, the recording, the manifest’s cache directory, and stopping the bot it spawned.

  • use_cache –

    The params field of the same name.

    Deprecated since version 1.9.0: Use params instead. Will be removed in 2.0.0.

  • default_timeout_ms –

    The params field of the same name.

    Deprecated since version 1.9.0: Use params instead. Will be removed in 2.0.0.