suite
Multi-bot eval suite runner.
An EvalManifest lists bots to spawn and the scenarios to run
against each, scripted and simulated alike. An EvalSuite spawns
each bot with its eval transport on its own port and drives it with the
harness in a subprocess, several at a time. pipecat eval suite is the
CLI in front of it; the release evals are a manifest plus that command.
Manifest format (YAML):
concurrency: 4
repeat: 1 # run each (bot, scenario) N times
runs_dir: test-runs # logs + recordings go to <runs_dir>/<timestamp>/
record: false # record conversation audio
cache_dir: null # optional
scenarios_dir: scenarios # resolved relative to this manifest file
# {python}=interpreter (default sys.executable), {bot}=bot path,
# {port}=assigned per run by the suite runner
spawn: "{python} {bot} -t eval --port {port}"
suite:
- bot: examples/voice/voice-cartesia.py
scenarios: [simple_math, greeting]
- bot: examples/voice/voice-openai.py
scenarios: [simple_math, interruption]
- bot: examples/voice/voice-groq.py
concurrency: 2 # this bot's own cap
scenarios: [simple_math, interruption]
- bot: examples/vision/vision-openai.py
runner_body:
path: scenarios/vision-cat.yaml # passed to the bot as --runner-body
scenarios: [vision_describe]
- bot: examples/turns/filter-incomplete-turns.py
name: openai/gpt-4o-mini # this entry's label, the bot path by default
runner_body:
data: {model: gpt-4o-mini} # written to a file for the bot
scenarios: [turn_completion]
- bot: examples/turns/filter-incomplete-turns.py
name: groq/llama-3.3-70b
runner_body:
data: {model: llama-3.3-70b}
scenarios: [turn_completion]
- bot: examples/flows/restaurant_reservation.py
scenarios: [book_table] # a simulation: its file has a persona
A scenarios: entry names a scenario file, and the file contributes one run
per scenario it holds, scripted or a simulation as each says (see
EvalScenarioFile), named
<file name>/<scenario name>. A name resolves under scenarios_dir with
.yaml added and may carry a folder, as scripted/greeting; a name ending
in .yaml is a path relative to the manifest instead. A simulation runs as
many times as its runs: says, and every run must pass.
An optional runner_body: supplies runner-args data the bot would normally
receive in a /start request body (e.g. a vision bot’s image path), passed
to it as --runner-body. It holds either path:, a YAML or JSON file
resolved relative to the manifest, or data:, the body itself as a mapping,
which the suite writes to a file among the run’s logs. A bot given a file runs
with the file’s directory as its working directory, so relative paths inside
the body (like an image) resolve next to the file; a body that holds such
paths belongs in a file for that reason.
Deprecated since version 1.11.0: Use runner_body: {path: <file>} instead of a bare runner_body: <file>.
Will be removed in 2.0.0.
An entry’s name: is its label: what the display, the -p filter, the
results records and the artifact file names use, and what an entry’s own
concurrency: is keyed on. It defaults to the bot: path, so it is only
needed when several entries share a bot, as when sweeping a model with
runner_body:; two entries may not run the same scenario under the same
label.
concurrency is how many runs execute at once. The suite keeps that many
going, taking the next run from the first entry in manifest order that still
has one, so an entry’s scenarios finish together and no slot waits while any
entry still has runs. An entry whose provider rate-limits sets its own
concurrency:, and never has more than that many runs in flight.
Manifest-relative paths (bot/bots_dir, scenarios_dir,
runs_dir) resolve relative to the manifest file, so a manifest is portable;
the same values passed as CLI overrides resolve against the working directory.
repeat (or --repeat) runs every (bot, scenario) pair N times, which is how
a flaky behavior gets measured rather than sampled: a bot that passes a scenario
half the time looks identical to a reliable one in a single pass. Attempts run
attempt-major, every entry’s first attempt before any entry’s second, and carry
an EvalRun.attempt number that joins their artifact filenames, so no
attempt overwrites another’s logs.
- pipecat.evals.suite.capture_pipeline_logs(logs_dir: Path, prefix: str, *, name: str, enabled: bool) Iterator[None][source]
Capture the harness’s logs for one run into a single
<prefix>.debug.log.The logs are buffered in memory and written on exit, one section per pipeline. The run is tagged with
prefixand the sink filters on it, so concurrent runs never mix. Writes nothing unlessenabled.- Parameters:
logs_dir – Directory the
<prefix>.debug.logis written to.prefix – Filename stem; also the
eval_runid the sink filters on.name – Human test name shown in each section heading.
enabled – When False, do nothing and write no file.
- class pipecat.evals.suite.EvalRun(bot: str, scenario: str, scenario_path: Path, name: str | None = None, loaded: EvalScriptScenario | EvalSimulationScenario | None = None, bot_path: Path | None = None, bot_url: str | None = None, runner_body_path: Path | None = None, runner_body: dict | None = None, concurrency: int | None = None, kind: EvalKind = EvalKind.SCRIPT, attempts: int = 1, sweep: bool = False, attempt: int = 1, status: str = 'pending', stopping: bool = False, result: EvalScriptResult | EvalSimulationResult | None = None, error: str | None = None, started_at: float | None = None, duration_ms: int | None = None)[source]
Bases:
objectMutable per-(bot, scenario) state, updated in place so a live display can read it.
- Parameters:
bot – The manifest’s
bot:path (suite) or the bot URL (run).name – The manifest entry’s
name:, orNonewhen it has none;labelis what the display, the filters and the results use.scenario – Display name (the scenario or simulation, without
.yaml).scenario_path – Path to the scenario or simulation file.
kind –
script(played byEvalScriptSession) orsimulation(EvalSimulationSession).attempts – How many times this (bot, scenario) pair runs: the manifest’s
repeat, or a simulation’s ownruns. Above 1, each attempt’s artifacts carry its number.sweep – Whether the attempts come from a
repeat(a measurement: the suite reports a rate and a failure is data) rather than from a simulation’sruns(a requirement: every attempt must pass).bot_path – The bot to spawn (suite);
Nonewhen connecting tobot_url.bot_url – Connect here instead of spawning (used by
pipecat eval run).runner_body_path – Optional
--runner-bodyfile for the bot’s runner args.runner_body – The bot’s runner-args body given inline, written to a file for the bot when it is spawned;
Nonewhen there is none or it comes fromrunner_body_path.concurrency – How many runs of this bot’s entry may be in flight at once, when its manifest entry says;
Noneis as many as the suite runs.attempt – 1-based attempt number when the suite repeats (see
EvalManifest.repeat); always 1 for a single pass.status –
pending,running, ordone.stopping – True while the bot is being stopped after the run is done. The run still counts as in flight, so the next one waits.
result – The outcome, once the run is done.
error – Spawn/connection error message, if the run failed before producing a result.
started_at – Monotonic start time, for the live elapsed counter.
duration_ms – Wall-clock time the run took, in milliseconds.
loaded – The scenario, when the run was built from a loaded file;
load()returns it, or reads the file when it isNone.
- loaded: EvalScriptScenario | EvalSimulationScenario | None = None
- result: EvalScriptResult | EvalSimulationResult | None = None
- load() EvalScriptScenario | EvalSimulationScenario[source]
The run’s scenario: as loaded when the run was built, else read from its file.
- Raises:
ValueError – If the file is invalid.
KeyError – If the file holds no scenario of this name.
FileNotFoundError – If the file doesn’t exist.
- class pipecat.evals.suite.EvalManifest(runs: list[EvalRun], spawn: str, python: str, concurrency: int, repeat: int, base_port: int, runs_dir: Path | None, record: bool, cache_dir: str | None)[source]
Bases:
objectA parsed eval-suite manifest.
- Parameters:
runs – The (bot, scenario) and (bot, simulation) runs to execute.
spawn – Spawn command template (
{python}/{bot}/{port}substituted).python – Interpreter used to spawn each bot.
concurrency – How many runs to execute at once.
repeat – How many times to run each (bot, scenario) pair. Attempts run attempt-major rather than grouped per bot, so every bot meets the same machine conditions in the same stretch of the sweep and a transient slowdown shows up as a band across all of them instead of a regression in whichever bot happened to be running.
base_port – First port to assign; each run gets
base_port + index, so the reserved range widens withrepeat(bots x scenarios x repeat ports frombase_portup).runs_dir – Base for run output (a
<name>/subdir is added), orNone.record – Whether to record conversation audio.
cache_dir – Directory for cached synthesized user audio, or
None.
- classmethod load(path: str | Path, *, bots_dir: str | Path | None = None, scenarios_dir: str | Path | None = None, runs_dir: str | Path | None = None, spawn: str | None = None, python: str | None = None, concurrency: int | None = None, repeat: int | None = None, base_port: int | None = None, record: bool | None = None, cache_dir: str | None = None) EvalManifest[source]
Parse a manifest YAML into an
EvalManifest.A keyword that is not
Noneoverrides the manifest’s value, so the CLI wins. Manifest paths resolve against the manifest’s directory, overrides against the working directory.- Parameters:
path – Path to the manifest YAML.
bots_dir – Override for the manifest’s
bots_dir(bot paths are relative to it).scenarios_dir – Override for the manifest’s
scenarios_dir.runs_dir – Override for the manifest’s
runs_dir(base for run output).spawn – Override for the spawn command template.
python – Override for the interpreter used to spawn bots.
concurrency – Override for how many runs execute at once.
repeat – Override for how many times each (bot, scenario) pair runs. Set here or in the manifest, it also replaces each simulation’s own
runs, a repeat of 1 included.base_port – Override for the first port assigned.
record – Override for whether to record conversation audio.
cache_dir – Override for the synthesized-audio cache directory.
- Returns:
The parsed
EvalManifest.
- class pipecat.evals.suite.EvalSuite(manifest: EvalManifest)[source]
Bases:
BaseObjectRuns the (bot, scenario) runs of an
EvalManifest, spawning each bot.Each bot gets its eval transport on its own port and is driven by the harness in a subprocess, several at a time up to the manifest’s
concurrency. The runs are updated in place as they go, so a live display can read their progress.Event handlers available:
on_update: Called with an
EvalRunwhenever that run changes status. Runs are mutated in place, so handlers run synchronously and must return promptly; they see the run as the change left it rather than however it has moved on since.
Example:
manifest = EvalManifest.load("manifest.yaml") suite = EvalSuite(manifest) suite.filter(pattern="voice") @suite.event_handler("on_update") async def on_update(suite, run): print(run.scenario, run.status) await suite.run(Path("logs"))
- __init__(manifest: EvalManifest)[source]
Initialize the suite from a parsed manifest.
- Parameters:
manifest – The parsed
EvalManifest; its runs become the suite’s working set (narrowed byfilter(), executed byrun()).
- filter(*, pattern: str | None = None, scenario: str | None = None, kind: EvalKind | None = None) list[EvalRun][source]
Keep only the runs matching a bot substring, a scenario name, and/or a kind.
- Parameters:
pattern – Keep only runs whose name or bot path contains this substring.
scenario – Keep only runs of this scenario: its full
<file>/<scenario>name, or either half of it, so a file’s name selects every scenario it holds.kind – Keep only runs of this kind.
- Returns:
The matching runs, in their original order.
- async run(logs_dir: Path, *, record_dir: Path | None = None, results_path: Path | None = None, on_update: Callable[[EvalRun], None] | None = None, debug: bool = False, params: EvalSessionParams | None = None, use_cache: bool | None = None, default_timeout_ms: int | None = None) None[source]
Run all of the suite’s runs, in place, with the manifest’s concurrency.
Each run gets its own port (
base_port + index).concurrencyworkers each take the next run and run it until none is left. Runs are taken in manifest order, every entry’s first attempt before any entry’s second, so an entry’s scenarios finish together and no worker waits while any entry still has runs. An entry with aconcurrency:of its own never has more than that many runs in flight, across its attempts; a worker that finds it full takes the next entry’s run.- Parameters:
logs_dir – Directory for per-run logs.
record_dir – Directory for per-run conversation recordings, or
None.results_path – JSONL file to append one record per finished run, or
Noneto write none. Each line is flushed as its run completes, so an interrupted sweep keeps everything already finished.on_update –
Called whenever a run changes status, for live display.
Deprecated since version 1.9.0: Use the
on_updateevent handler instead. Will be removed in 2.0.0.debug – When True, save each run’s combined
<run>.debug.log.params – How each run behaves;
Nonefor the defaults. The suite sets what it owns on each run’s copy: the connect timeout, the recording, the manifest’s cache directory, and stopping the bot it spawned.use_cache –
The
paramsfield of the same name.Deprecated since version 1.9.0: Use
paramsinstead. Will be removed in 2.0.0.default_timeout_ms –
The
paramsfield of the same name.Deprecated since version 1.9.0: Use
paramsinstead. Will be removed in 2.0.0.