simulation
Simulated scenario file format for Pipecat behavioral evaluations.
A simulation describes a caller rather than a script: who they are, what
they want, and how the outcome is judged. A persona LLM holds the
conversation with the bot on its own. A scenario with a persona: is a
simulation; it lives in a scenario file’s scenarios: list
(pipecat.evals.scenario) and a manifest lists the file like any other.
Example:
name: capital_curious
judge: !include judge_text.yaml
scenarios:
- name: capital_curious
persona: |
A curious, polite traveler who asks one thing at a time.
goal: "Find out what the capital of Germany is, then say goodbye."
success: "the bot told the caller that the capital of Germany is Berlin"
metrics:
- name: politeness
criterion: "the bot stayed courteous throughout"
min_score: 1
max_turns: 10
Fields:
persona,goalthe caller’s character and what they are trying to accomplish; both go into the persona LLM’s instructions (see
pipecat.evals.persona).simulatorthe persona LLM:
service(ollama) or afactory,model, and the optionalendpoint/extrathe judge config also takes. Omitted, the persona runs on the same local Ollama model as the default judge. The model must support function calling: the persona ends the call by calling itsend_calltool.user,judgethe blocks scenarios use (see
pipecat.evals.script):user.modalityanduser.speechdecide whether the persona’s turns reach the bot as synthesized speech or as text;judge.modality,judge.transcriptionandjudge.evaldecide whether the bot speaks and which LLM judges the outcome. Audio modality needs auser.speechblock, since every persona turn is synthesized.successwhat counts as the bot having done its job, judged over the whole conversation; prose, as long as it needs to be. It is the bot’s side of the
goal: usually that the caller got what they asked for, but where the right outcome is to refuse, to qualify, or to escalate, it says so. The run succeeded if the judge says yes. The judge sees the bot’s tool calls (name and arguments), not their results. Whether the bot made a call at all is afunction_callsmetric, no judge needed; if a reply must match backend data, write the expected value into the criterion (“the reply says the appointment is on Tuesday September fifteenth”) and keep the mocks deterministic so it stays true across runs.metricsjudged quality criteria, each with
name,criterion, and an optionalmin_scorein 0..1. A criterion says what every reply of the bot should be; the judge decides it for each bot turn, in the light of the conversation before it and the tool calls the bot had made by then, with a yes or a no, never a partial score. The metric’s score is the share of turns that got a yes: 0.80 is four replies in five. A turn the judge leaves out counts as a no, recorded as a verdict ofnoneso a sweep can tell judge trouble from bot trouble, and a run with no bot turn has no score and passes. A metric with amin_scorefails the run when its score is below it; one without is reported and never fails anything. Something the bot must do once, read the order back, belongs insuccess, not here.A metric can measure instead of judge:
measurenames one ofSIMULATION_MEASURESandmin_value/max_value(at least one) bound it; thenamedefaults to the measure. The harness computes the value from the run, no judge involved, and the metric scores 1 inside the range and 0 outside, which fails the run.turnsis the persona’s turns,durationthe conversation’s seconds from its first line to the hang-up,wordsthe longest bot reply in words, andlatencythe slowest reply in seconds: from the persona’s send to the reply’s first token in text mode, from the bot noticing the persona stop to its first spoken sentence in audio mode, which a failure’s reason spells out, since the two are not comparable. The per-reply measures bound every reply.function_callstakes acalls:list instead of a range: the calls the bot should make, each a name or anamewithargs(a subset of the call’s arguments), in any order. Every listed call must have happened and any call not listed fails it, socalls: []says the bot must call nothing, the check for a caller who should be turned down. A call the bot cancelled did not happen.max_turns,max_duration_s,max_silence_sbackstops on the persona’s turns (default 20), on the run’s wall clock (default 300 s), and on a lull in which neither side does anything (default 30 s), so a bot that never greets, or stops answering, ends the run as
silenceinstead of running out the clock. A run they end has not succeeded. A failure of the harness’s own pipeline, the persona LLM first among them, ends the run at once as an error.runshow many times the suite runs the simulation (default 1). Every run must pass: a persona does not say the same thing twice, so one run is an anecdote and three are a check.
- class pipecat.evals.simulation.EvalSimulationMetric(name: str, criterion: str | None = None, min_score: float | None = None, measure: str | None = None, min_value: float | None = None, max_value: float | None = None, calls: list[EvalFunctionCall] | None = None)[source]
Bases:
objectA quality metric: a judged criterion, or a measure with a range or a call list.
- Parameters:
name – The metric’s name in the results.
criterion – What the judge decides on each bot turn;
Nonefor a measured metric.min_score – The share of the bot’s turns the judge must answer yes for, in 0..1, below which a judged metric fails the run;
Nonereports the score without gating.measure – One of
SIMULATION_MEASURES;Nonefor a judged metric.min_value – The measured value’s lower bound, inclusive, or
None.max_value – The measured value’s upper bound, inclusive, or
None.calls – For
function_calls, the calls the bot should make: each a name, or a name withargsas a subset of the call’s arguments; an empty list means none.Nonefor every other metric.
- calls: list[EvalFunctionCall] | None = None
- class pipecat.evals.simulation.EvalSimulationScenario(name: str, persona: str, goal: str, success: str, simulator: dict = <factory>, metrics: list[EvalSimulationMetric] = <factory>, judge: dict = <factory>, bot_audio: bool = False, transcriber: dict | None = None, user_audio: bool = False, user_speech: dict | None = None, max_turns: int = 20, max_duration_s: float = 300.0, max_silence_s: float = 30.0, runs: int = 1, trigger_disconnect: bool = False, source_path: Path | None = None)[source]
Bases:
objectA parsed simulation file.
- Parameters:
name – The simulation name (from
name:).persona – Who the caller is, as free text for the persona LLM.
goal – What the caller wants from the call.
simulator – The persona LLM config (
service,model, optionalendpoint/extra), the same shape asjudge.eval; empty for the default local model.success – What counts as the bot having done its job, for the judge.
metrics – The judged quality criteria.
judge – Judge LLM config, as for a scenario.
bot_audio – Whether the bot speaks (
judge.modality: audio).transcriber – STT config for the bot’s audio in audio modality, else None.
user_audio – Whether the persona’s turns reach the bot as speech (
user.modality: audio).user_speech – TTS config the persona’s turns are synthesized with in audio modality, else None.
max_turns – Cap on the persona’s turns.
max_duration_s – Cap on the run’s wall clock, in seconds.
max_silence_s – Cap on a lull with no event from either side, in seconds.
runs – How many times the suite runs the simulation; every run must pass.
trigger_disconnect – Whether the harness fires the bot’s
on_client_disconnectedhandler when the connection ends.source_path – Path the simulation was loaded from, for error messages.
- classmethod load(path: str | Path) EvalSimulationScenario[source]
Parse a YAML file holding a simulation’s own keys at its top level.
Deprecated since version 1.11.0: Use
load()instead. Will be removed in 2.0.0.- Parameters:
path – Path to a YAML file with the simulation schema.
- Returns:
The parsed simulation.
- pipecat.evals.simulation.describe_simulation(simulation: EvalSimulationScenario, *, color: bool = False) str[source]
Three-line summary of a simulation’s user config, judge config, and goal, for pre-run logs.
The lines of
describe_config(), theuserline also naming the persona LLM that plays the user and the run’s caps, then the caller’s goal, e.g.:user -> modality: text | persona: ollama/gemma4:12b | max_turns: 8 | max_duration_s: 120 | max_silence_s: 30 judge -> modality: text | eval: ollama/gemma4:12b goal -> Book a table for two at 6 PM, then end the call.
- Parameters:
simulation – The parsed simulation to summarize.
color – When True, ANSI-color the keywords as
describe_configdoes.
- Returns:
The summary, one line per section.