simulation

Simulated scenario file format for Pipecat behavioral evaluations.

A simulation describes a caller rather than a script: who they are, what they want, and how the outcome is judged. A persona LLM holds the conversation with the bot on its own. A scenario with a persona: is a simulation; it lives in a scenario file’s scenarios: list (pipecat.evals.scenario) and a manifest lists the file like any other. Example:

name: capital_curious
judge: !include judge_text.yaml
scenarios:
  - name: capital_curious
    persona: |
      A curious, polite traveler who asks one thing at a time.
    goal: "Find out what the capital of Germany is, then say goodbye."
    success: "the bot told the caller that the capital of Germany is Berlin"
    metrics:
      - name: politeness
        criterion: "the bot stayed courteous throughout"
        min_score: 1
    max_turns: 10

Fields:

persona, goal

the caller’s character and what they are trying to accomplish; both go into the persona LLM’s instructions (see pipecat.evals.persona).

simulator

the persona LLM: service (ollama) or a factory, model, and the optional endpoint / extra the judge config also takes. Omitted, the persona runs on the same local Ollama model as the default judge. The model must support function calling: the persona ends the call by calling its end_call tool.

user, judge

the blocks scenarios use (see pipecat.evals.script): user.modality and user.speech decide whether the persona’s turns reach the bot as synthesized speech or as text; judge.modality, judge.transcription and judge.eval decide whether the bot speaks and which LLM judges the outcome. Audio modality needs a user.speech block, since every persona turn is synthesized.

success

what counts as the bot having done its job, judged over the whole conversation; prose, as long as it needs to be. It is the bot’s side of the goal: usually that the caller got what they asked for, but where the right outcome is to refuse, to qualify, or to escalate, it says so. The run succeeded if the judge says yes. The judge sees the bot’s tool calls (name and arguments), not their results. Whether the bot made a call at all is a function_calls metric, no judge needed; if a reply must match backend data, write the expected value into the criterion (“the reply says the appointment is on Tuesday September fifteenth”) and keep the mocks deterministic so it stays true across runs.

metrics

judged quality criteria, each with name, criterion, and an optional min_score in 0..1. A criterion says what every reply of the bot should be; the judge decides it for each bot turn, in the light of the conversation before it and the tool calls the bot had made by then, with a yes or a no, never a partial score. The metric’s score is the share of turns that got a yes: 0.80 is four replies in five. A turn the judge leaves out counts as a no, recorded as a verdict of none so a sweep can tell judge trouble from bot trouble, and a run with no bot turn has no score and passes. A metric with a min_score fails the run when its score is below it; one without is reported and never fails anything. Something the bot must do once, read the order back, belongs in success, not here.

A metric can measure instead of judge: measure names one of SIMULATION_MEASURES and min_value / max_value (at least one) bound it; the name defaults to the measure. The harness computes the value from the run, no judge involved, and the metric scores 1 inside the range and 0 outside, which fails the run. turns is the persona’s turns, duration the conversation’s seconds from its first line to the hang-up, words the longest bot reply in words, and latency the slowest reply in seconds: from the persona’s send to the reply’s first token in text mode, from the bot noticing the persona stop to its first spoken sentence in audio mode, which a failure’s reason spells out, since the two are not comparable. The per-reply measures bound every reply. function_calls takes a calls: list instead of a range: the calls the bot should make, each a name or a name with args (a subset of the call’s arguments), in any order. Every listed call must have happened and any call not listed fails it, so calls: [] says the bot must call nothing, the check for a caller who should be turned down. A call the bot cancelled did not happen.

max_turns, max_duration_s, max_silence_s

backstops on the persona’s turns (default 20), on the run’s wall clock (default 300 s), and on a lull in which neither side does anything (default 30 s), so a bot that never greets, or stops answering, ends the run as silence instead of running out the clock. A run they end has not succeeded. A failure of the harness’s own pipeline, the persona LLM first among them, ends the run at once as an error.

runs

how many times the suite runs the simulation (default 1). Every run must pass: a persona does not say the same thing twice, so one run is an anecdote and three are a check.

class pipecat.evals.simulation.EvalSimulationMetric(name: str, criterion: str | None = None, min_score: float | None = None, measure: str | None = None, min_value: float | None = None, max_value: float | None = None, calls: list[EvalFunctionCall] | None = None)[source]

Bases: object

A quality metric: a judged criterion, or a measure with a range or a call list.

Parameters:
  • name – The metric’s name in the results.

  • criterion – What the judge decides on each bot turn; None for a measured metric.

  • min_score – The share of the bot’s turns the judge must answer yes for, in 0..1, below which a judged metric fails the run; None reports the score without gating.

  • measure – One of SIMULATION_MEASURES; None for a judged metric.

  • min_value – The measured value’s lower bound, inclusive, or None.

  • max_value – The measured value’s upper bound, inclusive, or None.

  • calls – For function_calls, the calls the bot should make: each a name, or a name with args as a subset of the call’s arguments; an empty list means none. None for every other metric.

criterion: str | None = None
min_score: float | None = None
measure: str | None = None
min_value: float | None = None
max_value: float | None = None
calls: list[EvalFunctionCall] | None = None
class pipecat.evals.simulation.EvalSimulationScenario(name: str, persona: str, goal: str, success: str, simulator: dict = <factory>, metrics: list[EvalSimulationMetric] = <factory>, judge: dict = <factory>, bot_audio: bool = False, transcriber: dict | None = None, user_audio: bool = False, user_speech: dict | None = None, max_turns: int = 20, max_duration_s: float = 300.0, max_silence_s: float = 30.0, runs: int = 1, trigger_disconnect: bool = False, source_path: Path | None = None)[source]

Bases: object

A parsed simulation file.

Parameters:
  • name – The simulation name (from name:).

  • persona – Who the caller is, as free text for the persona LLM.

  • goal – What the caller wants from the call.

  • simulator – The persona LLM config (service, model, optional endpoint / extra), the same shape as judge.eval; empty for the default local model.

  • success – What counts as the bot having done its job, for the judge.

  • metrics – The judged quality criteria.

  • judge – Judge LLM config, as for a scenario.

  • bot_audio – Whether the bot speaks (judge.modality: audio).

  • transcriber – STT config for the bot’s audio in audio modality, else None.

  • user_audio – Whether the persona’s turns reach the bot as speech (user.modality: audio).

  • user_speech – TTS config the persona’s turns are synthesized with in audio modality, else None.

  • max_turns – Cap on the persona’s turns.

  • max_duration_s – Cap on the run’s wall clock, in seconds.

  • max_silence_s – Cap on a lull with no event from either side, in seconds.

  • runs – How many times the suite runs the simulation; every run must pass.

  • trigger_disconnect – Whether the harness fires the bot’s on_client_disconnected handler when the connection ends.

  • source_path – Path the simulation was loaded from, for error messages.

bot_audio: bool = False
transcriber: dict | None = None
user_audio: bool = False
user_speech: dict | None = None
max_turns: int = 20
max_duration_s: float = 300.0
max_silence_s: float = 30.0
runs: int = 1
trigger_disconnect: bool = False
source_path: Path | None = None
classmethod load(path: str | Path) → EvalSimulationScenario[source]

Parse a YAML file holding a simulation’s own keys at its top level.

Deprecated since version 1.11.0: Use load() instead. Will be removed in 2.0.0.

Parameters:

path – Path to a YAML file with the simulation schema.

Returns:

The parsed simulation.

pipecat.evals.simulation.describe_simulation(simulation: EvalSimulationScenario, *, color: bool = False) → str[source]

Three-line summary of a simulation’s user config, judge config, and goal, for pre-run logs.

The lines of describe_config(), the user line also naming the persona LLM that plays the user and the run’s caps, then the caller’s goal, e.g.:

user  -> modality: text | persona: ollama/gemma4:12b | max_turns: 8 | max_duration_s: 120 | max_silence_s: 30
judge -> modality: text | eval: ollama/gemma4:12b
goal  -> Book a table for two at 6 PM, then end the call.
Parameters:
  • simulation – The parsed simulation to summarize.

  • color – When True, ANSI-color the keywords as describe_config does.

Returns:

The summary, one line per section.