script
Scripted scenario file format for Pipecat behavioral evaluations.
A scripted scenario describes a conversation and the semantic events expected
to flow back from the bot. It lives in a scenario file’s scenarios: list
(pipecat.evals.scenario). Simple example:
name: simple_user_input
scenarios:
- name: simple_user_input
turns:
- user: "hello world"
expect:
- event: user_started_speaking
- event: user_transcription
text_contains: "hello world"
The harness plays each turn and checks the events the bot emits back, in
order. A scenario with a persona: instead of turns: is the other kind,
a simulation, where an LLM plays the user (pipecat.evals.simulation).
The keys below are a scenario’s; the file-level ones (user:, judge:,
context:, stop_on_failure:) may also sit at the top of the file as
defaults for every scenario in it.
Event names are the friendly names the harness maps RTVI server messages onto:
user_started_speaking, user_stopped_speaking, vad_user_started_speaking,
vad_user_stopped_speaking, user_transcription, bot_started_speaking,
bot_stopped_speaking, llm_started, response, llm_response,
llm_marker, tts_response, function_call, function_call_stopped,
image. The vad_* events are the raw
VAD signal, useful as a timing anchor when a turn-detection strategy gates or defers the
turn-level user_stopped_speaking (e.g. filtering incomplete turns).
image is an image the bot output, such as a generated one; the bot reports
images only to a scenario that expects one:
turns:
- expect:
- event: image
within_ms: 60000
llm_marker is the sideband marker the bot’s LLM emitted in a response, such
as the turn-completion markers of filter_incomplete_user_turns. It arrives
when the response ends. marker: says which one: complete (the turn was
finished and the bot answers), short (the user was cut off and the bot
waits), long (the user asked for time), or incomplete (either of the
last two). A bare llm_marker asserts only that the bot read a response.
Markers never reach clients by default; a scenario that asserts on one asks the
bot to report them:
turns:
- user: "Let me think about it, hmmm"
expect:
- event: llm_marker
marker: incomplete # the bot held the turn open
- user: "I think I'd go to Japan."
expect:
- event: llm_marker
marker: complete # ... and answered this one
- event: response
eval: "engages with the user's answer about Japan"
The event also carries the response’s raw text, as the LLM produced it before
the bot held anything back, so a scenario can check how well the LLM follows
the protocol: marker_first (nothing before the marker), markers (how
many markers the text holds), and text_after (whether text follows the
first marker, which a complete turn should have and an incomplete one should
not):
- user: "I'd go to Japan because"
expect:
- event: llm_marker
marker: short
marker_first: true
markers: 1
text_after: false
A marker the LLM lets slip into its reply reaches the user, so the reply’s
own text is worth checking too: text_excludes fails when the text holds
the given string, the mirror of text_contains:
- user: "What is the capital of Germany?"
expect:
- event: llm_response
text_contains: Berlin
text_excludes: "●"
The bot’s reply can be asserted three ways:
responsethe transcription of the bot’s actual synthesized audio (a local STT — Moonshine or Whisper — run by the harness) in audio modality, or the LLM text in text modality. The real end-to-end check — prefer this.
llm_responsethe LLM’s text output (
bot-llm-text). Available in both modalities.tts_responsethe text the TTS reports speaking (
bot-tts-text, with word timing). Audio modality only.
Supported expectation fields (per event):
event: <name>required — event type name
within_ms: <int>latency budget from the most recent anchor (optional; defaults to 60s when omitted)
text_contains: <str>substring check on the event’s text content, ignoring whitespace differences
text_excludes: <str>the reverse: the event’s text content must not hold this substring
marker: <str>for
llm_marker— the marker’s meaning:complete,short,long, orincompletefor either of the last twomarker_first: <bool>,markers: <int>,text_after: <bool>for
llm_marker— checks on the response’s raw text: whether the first marker comes before any text, how many markers the text holds, and whether text follows the first markercalls:for
function_call— the set of calls the turn should make, matched by name in any order; the expectation passes only when all are found:- event: function_call calls: - name: get_current_weather args: { location: San Francisco } - name: get_restaurant_recommendation
function_call_stoppedtakes the samecalls:shape, and itsargssay how the call ended — which is how a scenario tells work that was stopped from work that finished on its own:- event: function_call_stopped calls: - name: write_report args: { cancelled: true }
eval: <str>natural-language criterion the event’s text content must satisfy, evaluated by the judge (see
pipecat.evals.judge).On
function_callthe criterion is about the call instead: each callcalls:(or thename:/args:shorthand) matches is put to the judge by name and arguments, over the conversation so far, which is how a scenario checks whatargs:cannot match verbatim. Afunction_call_stoppedcarries no arguments, only how the call ended, so it takes noeval::- event: function_call calls: - name: submit_session_suggestion eval: "a session about OpenTelemetry tracing, submitted for Jennifer Smith"
absent: trueinvert the expectation: assert that NO event of this type arrives before the
within_msbudget expires (default 60s — setwithin_msexplicitly to keep the quiet-window wait short). Matches on event type only, so it cannot be combined withtext_contains,eval:, orcalls:. Aresponsethat continues the reply an earlier expectation matched is not a new one; only a reply the bot began after that match counts. Used for duplicate-output regressions:- event: response eval: "answers the question" - event: response absent: true within_ms: 30000
Instead of user:, a turn may press DTMF keys with dtmf: (the two are
mutually exclusive — you press keys or you talk):
turns:
- dtmf: "123#" # quote it: an unquoted # starts a YAML comment
expect:
- event: user_transcription
text_contains: "DTMF: 123#"
- event: response
eval: "confirms the entered digits"
Each character is sent as one InputDTMFFrame (0-9, *, #),
regardless of the scenario’s user/judge modality. A bot running a
DTMFAggregator accumulates them and flushes — on the # terminator or its
idle timeout — into a DTMF: ... transcription it reacts to.
A turn is sent once the bot has finished speaking, like a caller who waits for
the end of the sentence, so a reply the previous turn was satisfied with early
is never talked over. A turn may also include send_after: to schedule its
user/dtmf send
relative to a prior event (used for interruption / barge-in tests), or
image: (a path, relative to the scenario file) to register an image for the
turn — when a function-calling-video bot requests a user image, the eval
transport serves it. send_after with only delay_ms (no event) is a
pure time delay relative to the previous send — handy for pacing keypresses
across dtmf turns to exercise the aggregator’s idle-timeout flush.
In audio modality a turn may name a recording with audio: (a path, relative
to the scenario file) that is played to the bot in place of synthesizing its
user text: a real caller’s voice, or a clip that reproduces a bug. The file
is sent at its own sample rate, so it need not match the bot’s input rate, and
user still gives what the recording says, since that is what the judge and
text_contains see as the turn’s input:
turns:
- user: "What is the capital of Germany?"
audio: ../assets/capital_question.wav
expect:
- event: response
eval: "the response says the capital of Germany is Berlin"
expect: is optional; omit it for a turn that only sends input or only waits.
Top-level optional fields:
context:LLM messages the bot’s context should start from. When given, the harness sends them before driving turns (replacing the bot’s context); omit to leave the bot’s own context untouched.
stop_on_failure:whether the first failed turn ends the scenario (default true). A failure leaves the conversation in an unknown state, so the remaining turns usually just burn a timeout each. Set it false for a scenario that scores every turn independently — a benchmark that reports a per-turn pass rate needs all of its turns driven, not just the ones before the first miss:
stop_on_failure: false
Give those turns an explicit
within_ms: with the 60s default, a silent bot costs one full budget per remaining turn.user:how user turns are delivered:
user: modality: audio # audio | text (default text) speech: # needed unless every spoken turn has audio: service: kokoro # local TTS that synthesizes the user turns voice: af_heart # voices are language-specific language: en # optional; must match the voice speed: 1.0 # optional; Kokoro's rate, pauses included sample_rate: 16000 # optional # or, for any other TTS: factory: my_evals.voice (a callable # taking this mapping and returning a local or HTTP TTSService)
audiostreams synthesized user audio to the bot (exercising its STT for real);text(the default) sends RTVIsend-text. A scenario whose spoken turns all name anaudio:recording needs nospeech:block.judge:what the judge evaluates, and what decides the verdicts:
judge: modality: audio # audio | text (default text) eval: # what judges (default ollama) service: ollama model: gemma4:12b # or factory: my_evals.judge, a callable taking this mapping and # returning a BaseClassifier or an OpenAI-compatible LLM service # explainer: an LLM block giving the reasons behind the verdicts # (see pipecat.evals.judge) transcription: # required when modality is audio service: moonshine # STT for the bot's audio (or whisper, or a factory) model: small-streaming # optional language: en # optional; the language the bot speaks padding_secs: 0 # optional; silence padded around the # segment (default: 2)
audiomakes the bot speak and judges the transcription of its actual audio (tts_response);text(the default) skips TTS and judges the LLM text (llm_response), which is faster and silent.
Any value can be pulled from a separate file with !include, resolved
relative to the scenario file’s directory. This is handy for sharing the
judge: and user: blocks across scenarios:
user: !include user_audio.yaml
judge: !include judge_audio.yaml
- class pipecat.evals.script.EvalFunctionCall(name: str | None = None, args: dict | None = None)[source]
Bases:
objectOne expected function call within a
function_callexpectation.- Parameters:
name – The function name to match.
Nonematches any call (used by a barefunction_callexpectation that just asserts a call happened).args – Optional subset check on the call’s arguments (every listed key/value must be present; extra arguments are ignored).
- class pipecat.evals.script.EvalExpectation(event: str, within_ms: int | None = None, text_contains: str | None = None, text_excludes: str | None = None, calls: list[EvalFunctionCall] | None = None, eval: str | None = None, marker: str | None = None, marker_first: bool | None = None, markers: int | None = None, text_after: bool | None = None, absent: bool = False)[source]
Bases:
objectA single expected event in a scenario turn.
- Parameters:
event – Required — the semantic event name (e.g.
user_stopped_speaking).within_ms – Optional latency budget, measured from the turn’s user send — all of a turn’s expectations share that one anchor, so a stalled turn fails within a single budget rather than one per expectation. For audio turns the anchor is when the utterance was sent, not when it finishes streaming to the bot. Defaults to 60s when omitted, so timing isn’t asserted unless set explicitly.
text_contains – Optional substring check on the event’s text content (
llm_response.textoruser_transcription.transcript).text_excludes – Optional substring the event’s text content must not hold. Checked on the text the expectation matched; with
text_contains, on the reply accumulated up to the match.calls – For a
function_callevent, the set of calls expected in the turn. They are matched by name in any order and the expectation passes only when all of them are found. Built fromcalls:in the YAML, or from the singlename:/args:shorthand.eval – Optional natural-language criterion the event’s text content must satisfy. Evaluated by a judge LLM. Meaningful on the bot-generated text events (
response,llm_response, andtts_response) and on the function-call events, where it is about each matched call’s name and arguments rather than text.marker – For an
llm_markerevent, the meaning the marker must have: one ofMARKER_KINDS, whereincompleteacceptsshortorlong.marker_first – For an
llm_markerevent, whether the first marker in the response’s raw text must come before any text.markers – For an
llm_markerevent, how many markers the response’s raw text must hold.text_after – For an
llm_markerevent, whether text must (True) or must not (False) follow the first marker in the raw text.absent – When True, the expectation is inverted: it passes only when NO event of this type arrives before the
within_msbudget expires, and fails as soon as one does. Matches on event type only;text_contains,text_excludes,eval,callsand the marker checks are not allowed alongside it.
- calls: list[EvalFunctionCall] | None = None
- class pipecat.evals.script.EvalSendAfter(event: str | None, delay_ms: int)[source]
Bases:
objectScheduling for when a turn’s input (
userordtmf) is sent.When set on a
EvalScriptTurn, the harness waits foreventto have been seen (either earlier in the run or arriving now), then waits an additionaldelay_msbefore sending the turn’s input. Used for barge-in tests:send_after: {event: llm_started, delay_ms: 500}means “interrupt 500ms after the bot started responding.”eventis optional: a baresend_after: {delay_ms: 500}is a pure time delay with no event anchor (500ms after the previous turn’s send). Handy for pacing keypresses acrossdtmfturns, where there is no per-key event to anchor on.- Parameters:
event – Name of the event to schedule from, or
Nonefor a puredelay_mstime delay with no event anchor.delay_ms – Additional delay in milliseconds after the event was received (or, when
eventisNone, after the previous turn’s send).
- class pipecat.evals.script.EvalScriptTurn(user: str | None, expect: list[EvalExpectation] = <factory>, audio: str | None = None, dtmf: str | None = None, send_after: EvalSendAfter | None = None, image: str | None = None)[source]
Bases:
objectOne turn in a scenario.
A turn drives the bot one of three ways: the harness sends a
userutterance (the person speaks), it sends adtmfkeypress sequence (the person presses keys), or it is observation-only (neither field — useful for bot-first scenarios like opening greetings).useranddtmfare mutually exclusive: a turn is one or the other.A
userturn is spoken by the harness’s TTS unless the turn names anaudiofile to play instead.- Parameters:
user – Optional text the harness sends as the user’s turn — an RTVI
send-textin text modality, or synthesized speech (raw-audio) in audio modality. If absent, the turn just waits for and asserts on expected events.audio – Optional path to an audio file (resolved relative to the scenario file) played as this turn instead of synthesizing
user. Any formatsoundfilereads works (WAV, MP3, FLAC, OGG, …); multi-channel audio is downmixed to mono.useris required alongside it and is what the recording says — the judge reads it as the turn’s input andtext_containsmatches against it, neither of which can be recovered from the audio. Requiresuser.modality: audio, and is mutually exclusive withdtmf.dtmf – Optional DTMF keypad sequence the harness sends, one
InputDTMFFrameper character (e.g."123#"). Each character must be a validKeypadEntry(0-9,*,#). Mutually exclusive withuser. The keys are injected the same way regardless of the scenario’s user/judge modality; a bot with aDTMFAggregatorturns them into a transcription it reacts to. Quote the value in YAML (dtmf: "123#") — an unquoted#starts a comment.expect – Expected events, in the order they should arrive. Optional — omit it for a pure pacing/observation turn (e.g. a
dtmfturn that only presses keys, with the assertion on a later turn).send_after – Optional schedule for when the turn’s input should fire. Only meaningful when
userordtmfis set.image – Optional path to an image to register for this turn (resolved relative to the scenario file). When a function-calling-video bot requests a user image during the turn, the eval transport serves this one. Stays registered until a later turn provides a different image.
- send_after: EvalSendAfter | None = None
- class pipecat.evals.script.EvalTurn(**kwargs)[source]
Bases:
EvalScriptTurnDeprecated alias for
EvalScriptTurn.Deprecated since version 1.9.0: Use
EvalScriptTurninstead. Will be removed in 2.0.0.
- class pipecat.evals.script.EvalScriptScenario(name: str, turns: list[~pipecat.evals.script.EvalScriptTurn], context: list[dict] = <factory>, judge: dict = <factory>, bot_audio: bool = False, transcriber: dict | None = None, user_audio: bool = False, user_speech: dict | None = None, trigger_disconnect: bool = False, stop_on_failure: bool = True, source_path: ~pathlib.Path | None = None)[source]
Bases:
objectA parsed scenario file.
- Parameters:
name – The eval name (from
name:).turns – Ordered list of turns.
context – LLM messages the bot’s context should start from for this eval. When non-empty, the harness sends them as an
eval-contextclient message right after the bot-ready handshake (the eval serializer turns it into anLLMMessagesUpdateFrame, which replaces the context); bots without an LLM context aggregator ignore the frame. Omitted or empty (the default): the harness sends nothing and the bot keeps the context it set up itself.judge – Judge LLM configuration dict with keys
service,model, optionalendpoint, and an optionalextramapping forwarded to the model as top-level request parameters. Defaults to{"service": "ollama", "model": "gemma4:12b", "extra": {"reasoning_effort": "none"}}.bot_audio – Whether the bot produces speech, derived from
judge.modality. False (text, the default): the bot skips TTS — the harness configures skip-TTS at connect, so even an on-connect greeting is silent. True (audio): the bot speaks, and the judge evaluates the transcription of its actual audio.transcriber – Parsed from the
judge.transcription:block; the STT config (servicedefaults tomoonshine, plusmodeland an optionallanguagecode) used to transcribe the bot’s audio for theresponseevent (Nonein text modality). Setlanguagewhen the bot speaks a non-English language so the STT doesn’t default to English.user_audio – Whether the user’s turns reach the bot as speech, derived from
user.modality. False (text, the default): each turn is sent as an RTVIsend-text. True (audio): the harness streams RTVIraw-audio, exercising the bot’s STT for real.user_speech – Parsed from the
user.speech:block; the TTS config the harness synthesizes user turns with (Nonein text modality). Mapping withservice,voice, and optionalmodel/language/speed/sample_rate/api_key. Setlanguage(a code likezh) to synthesize non-English user turns.trigger_disconnect – Whether the harness fires the bot’s
on_client_disconnectedhandler when this scenario’s connection ends. Bots often cancel their pipeline there, so this is False by default to avoid that between scenarios; set True to exercise the bot’s disconnect path. Independent of--stop-bot, which tears the bot down viaeval-cancelregardless of the handler.stop_on_failure – Whether the first failed turn ends the scenario (default True). A failed turn leaves the conversation in an unknown state, so continuing usually costs one timeout per remaining turn. Set False for a scenario whose turns are scored independently, where the turns after a failure are still worth driving; each turn’s outcome is reported in
turns. This governs turn-to-turn progression only: within a turn, an expectation that times out still ends that turn’s matching, because a turn’s expectations share one deadline anchored at the send.source_path – Path the scenario was loaded from, for error messages.
- classmethod load(path: str | Path) EvalScriptScenario[source]
Parse a YAML file holding a scenario’s own keys at its top level.
Deprecated since version 1.11.0: Use
load()instead. Will be removed in 2.0.0.- Parameters:
path – Path to a YAML file with the scenario schema.
- Returns:
The parsed scenario.
- wants_response() bool[source]
Whether any expectation asserts on the transcription of the bot’s audio.
- required_report_level() str | None[source]
The function-call report level the scenario’s assertions need:
fullfor args,namefor names, elseNone.A call with an
eval:needsfulltoo: the judge is asked about the call’s arguments.
- needs_vad_events() bool[source]
Whether the scenario uses the raw VAD speaking events, which the bot emits only on request.
- class pipecat.evals.script.EvalScenario(**kwargs)[source]
Bases:
EvalScriptScenarioDeprecated alias for
EvalScriptScenario.Deprecated since version 1.9.0: Use
EvalScriptScenarioinstead. Will be removed in 2.0.0.