script

Scripted scenario file format for Pipecat behavioral evaluations.

A scripted scenario describes a conversation and the semantic events expected to flow back from the bot. It lives in a scenario file’s scenarios: list (pipecat.evals.scenario). Simple example:

name: simple_user_input
scenarios:
  - name: simple_user_input
    turns:
      - user: "hello world"
        expect:
          - event: user_started_speaking
          - event: user_transcription
            text_contains: "hello world"

The harness plays each turn and checks the events the bot emits back, in order. A scenario with a persona: instead of turns: is the other kind, a simulation, where an LLM plays the user (pipecat.evals.simulation). The keys below are a scenario’s; the file-level ones (user:, judge:, context:, stop_on_failure:) may also sit at the top of the file as defaults for every scenario in it.

Event names are the friendly names the harness maps RTVI server messages onto: user_started_speaking, user_stopped_speaking, vad_user_started_speaking, vad_user_stopped_speaking, user_transcription, bot_started_speaking, bot_stopped_speaking, llm_started, response, llm_response, llm_marker, tts_response, function_call, function_call_stopped, image. The vad_* events are the raw VAD signal, useful as a timing anchor when a turn-detection strategy gates or defers the turn-level user_stopped_speaking (e.g. filtering incomplete turns).

image is an image the bot output, such as a generated one; the bot reports images only to a scenario that expects one:

turns:
  - expect:
      - event: image
        within_ms: 60000

llm_marker is the sideband marker the bot’s LLM emitted in a response, such as the turn-completion markers of filter_incomplete_user_turns. It arrives when the response ends. marker: says which one: complete (the turn was finished and the bot answers), short (the user was cut off and the bot waits), long (the user asked for time), or incomplete (either of the last two). A bare llm_marker asserts only that the bot read a response. Markers never reach clients by default; a scenario that asserts on one asks the bot to report them:

turns:
  - user: "Let me think about it, hmmm"
    expect:
      - event: llm_marker
        marker: incomplete        # the bot held the turn open
  - user: "I think I'd go to Japan."
    expect:
      - event: llm_marker
        marker: complete          # ... and answered this one
      - event: response
        eval: "engages with the user's answer about Japan"

The event also carries the response’s raw text, as the LLM produced it before the bot held anything back, so a scenario can check how well the LLM follows the protocol: marker_first (nothing before the marker), markers (how many markers the text holds), and text_after (whether text follows the first marker, which a complete turn should have and an incomplete one should not):

- user: "I'd go to Japan because"
  expect:
    - event: llm_marker
      marker: short
      marker_first: true
      markers: 1
      text_after: false

A marker the LLM lets slip into its reply reaches the user, so the reply’s own text is worth checking too: text_excludes fails when the text holds the given string, the mirror of text_contains:

- user: "What is the capital of Germany?"
  expect:
    - event: llm_response
      text_contains: Berlin
      text_excludes: "●"

The bot’s reply can be asserted three ways:

response

the transcription of the bot’s actual synthesized audio (a local STT — Moonshine or Whisper — run by the harness) in audio modality, or the LLM text in text modality. The real end-to-end check — prefer this.

llm_response

the LLM’s text output (bot-llm-text). Available in both modalities.

tts_response

the text the TTS reports speaking (bot-tts-text, with word timing). Audio modality only.

Supported expectation fields (per event):

event: <name>

required — event type name

within_ms: <int>

latency budget from the most recent anchor (optional; defaults to 60s when omitted)

text_contains: <str>

substring check on the event’s text content, ignoring whitespace differences

text_excludes: <str>

the reverse: the event’s text content must not hold this substring

marker: <str>

for llm_marker — the marker’s meaning: complete, short, long, or incomplete for either of the last two

marker_first: <bool>, markers: <int>, text_after: <bool>

for llm_marker — checks on the response’s raw text: whether the first marker comes before any text, how many markers the text holds, and whether text follows the first marker

calls:

for function_call — the set of calls the turn should make, matched by name in any order; the expectation passes only when all are found:

- event: function_call
  calls:
    - name: get_current_weather
      args: { location: San Francisco }
    - name: get_restaurant_recommendation

function_call_stopped takes the same calls: shape, and its args say how the call ended — which is how a scenario tells work that was stopped from work that finished on its own:

- event: function_call_stopped
  calls:
    - name: write_report
      args: { cancelled: true }
eval: <str>

natural-language criterion the event’s text content must satisfy, evaluated by the judge (see pipecat.evals.judge).

On function_call the criterion is about the call instead: each call calls: (or the name:/args: shorthand) matches is put to the judge by name and arguments, over the conversation so far, which is how a scenario checks what args: cannot match verbatim. A function_call_stopped carries no arguments, only how the call ended, so it takes no eval::

- event: function_call
  calls:
    - name: submit_session_suggestion
  eval: "a session about OpenTelemetry tracing, submitted for Jennifer Smith"
absent: true

invert the expectation: assert that NO event of this type arrives before the within_ms budget expires (default 60s — set within_ms explicitly to keep the quiet-window wait short). Matches on event type only, so it cannot be combined with text_contains, eval:, or calls:. A response that continues the reply an earlier expectation matched is not a new one; only a reply the bot began after that match counts. Used for duplicate-output regressions:

- event: response
  eval: "answers the question"
- event: response
  absent: true
  within_ms: 30000

Instead of user:, a turn may press DTMF keys with dtmf: (the two are mutually exclusive — you press keys or you talk):

turns:
  - dtmf: "123#"            # quote it: an unquoted # starts a YAML comment
    expect:
      - event: user_transcription
        text_contains: "DTMF: 123#"
      - event: response
        eval: "confirms the entered digits"

Each character is sent as one InputDTMFFrame (0-9, *, #), regardless of the scenario’s user/judge modality. A bot running a DTMFAggregator accumulates them and flushes — on the # terminator or its idle timeout — into a DTMF: ... transcription it reacts to.

A turn is sent once the bot has finished speaking, like a caller who waits for the end of the sentence, so a reply the previous turn was satisfied with early is never talked over. A turn may also include send_after: to schedule its user/dtmf send relative to a prior event (used for interruption / barge-in tests), or image: (a path, relative to the scenario file) to register an image for the turn — when a function-calling-video bot requests a user image, the eval transport serves it. send_after with only delay_ms (no event) is a pure time delay relative to the previous send — handy for pacing keypresses across dtmf turns to exercise the aggregator’s idle-timeout flush.

In audio modality a turn may name a recording with audio: (a path, relative to the scenario file) that is played to the bot in place of synthesizing its user text: a real caller’s voice, or a clip that reproduces a bug. The file is sent at its own sample rate, so it need not match the bot’s input rate, and user still gives what the recording says, since that is what the judge and text_contains see as the turn’s input:

turns:
  - user: "What is the capital of Germany?"
    audio: ../assets/capital_question.wav
    expect:
      - event: response
        eval: "the response says the capital of Germany is Berlin"

expect: is optional; omit it for a turn that only sends input or only waits.

Top-level optional fields:

context:

LLM messages the bot’s context should start from. When given, the harness sends them before driving turns (replacing the bot’s context); omit to leave the bot’s own context untouched.

stop_on_failure:

whether the first failed turn ends the scenario (default true). A failure leaves the conversation in an unknown state, so the remaining turns usually just burn a timeout each. Set it false for a scenario that scores every turn independently — a benchmark that reports a per-turn pass rate needs all of its turns driven, not just the ones before the first miss:

stop_on_failure: false

Give those turns an explicit within_ms: with the 60s default, a silent bot costs one full budget per remaining turn.

user:

how user turns are delivered:

user:
  modality: audio          # audio | text (default text)
  speech:                  # needed unless every spoken turn has audio:
    service: kokoro        # local TTS that synthesizes the user turns
    voice: af_heart        # voices are language-specific
    language: en           # optional; must match the voice
    speed: 1.0             # optional; Kokoro's rate, pauses included
    sample_rate: 16000     # optional
    # or, for any other TTS: factory: my_evals.voice (a callable
    # taking this mapping and returning a local or HTTP TTSService)

audio streams synthesized user audio to the bot (exercising its STT for real); text (the default) sends RTVI send-text. A scenario whose spoken turns all name an audio: recording needs no speech: block.

judge:

what the judge evaluates, and what decides the verdicts:

judge:
  modality: audio          # audio | text (default text)
  eval:                    # what judges (default ollama)
    service: ollama
    model: gemma4:12b
    # or factory: my_evals.judge, a callable taking this mapping and
    # returning a BaseClassifier or an OpenAI-compatible LLM service
    # explainer: an LLM block giving the reasons behind the verdicts
    # (see pipecat.evals.judge)
  transcription:           # required when modality is audio
    service: moonshine     # STT for the bot's audio (or whisper, or a factory)
    model: small-streaming # optional
    language: en           # optional; the language the bot speaks
    padding_secs: 0        # optional; silence padded around the
                           # segment (default: 2)

audio makes the bot speak and judges the transcription of its actual audio (tts_response); text (the default) skips TTS and judges the LLM text (llm_response), which is faster and silent.

Any value can be pulled from a separate file with !include, resolved relative to the scenario file’s directory. This is handy for sharing the judge: and user: blocks across scenarios:

user: !include user_audio.yaml
judge: !include judge_audio.yaml
class pipecat.evals.script.EvalFunctionCall(name: str | None = None, args: dict | None = None)[source]

Bases: object

One expected function call within a function_call expectation.

Parameters:
  • name – The function name to match. None matches any call (used by a bare function_call expectation that just asserts a call happened).

  • args – Optional subset check on the call’s arguments (every listed key/value must be present; extra arguments are ignored).

name: str | None = None
args: dict | None = None
property signature: str

name(arg=value, ...).

Type:

A short label for the call

class pipecat.evals.script.EvalExpectation(event: str, within_ms: int | None = None, text_contains: str | None = None, text_excludes: str | None = None, calls: list[EvalFunctionCall] | None = None, eval: str | None = None, marker: str | None = None, marker_first: bool | None = None, markers: int | None = None, text_after: bool | None = None, absent: bool = False)[source]

Bases: object

A single expected event in a scenario turn.

Parameters:
  • event – Required — the semantic event name (e.g. user_stopped_speaking).

  • within_ms – Optional latency budget, measured from the turn’s user send — all of a turn’s expectations share that one anchor, so a stalled turn fails within a single budget rather than one per expectation. For audio turns the anchor is when the utterance was sent, not when it finishes streaming to the bot. Defaults to 60s when omitted, so timing isn’t asserted unless set explicitly.

  • text_contains – Optional substring check on the event’s text content (llm_response.text or user_transcription.transcript).

  • text_excludes – Optional substring the event’s text content must not hold. Checked on the text the expectation matched; with text_contains, on the reply accumulated up to the match.

  • calls – For a function_call event, the set of calls expected in the turn. They are matched by name in any order and the expectation passes only when all of them are found. Built from calls: in the YAML, or from the single name:/args: shorthand.

  • eval – Optional natural-language criterion the event’s text content must satisfy. Evaluated by a judge LLM. Meaningful on the bot-generated text events (response, llm_response, and tts_response) and on the function-call events, where it is about each matched call’s name and arguments rather than text.

  • marker – For an llm_marker event, the meaning the marker must have: one of MARKER_KINDS, where incomplete accepts short or long.

  • marker_first – For an llm_marker event, whether the first marker in the response’s raw text must come before any text.

  • markers – For an llm_marker event, how many markers the response’s raw text must hold.

  • text_after – For an llm_marker event, whether text must (True) or must not (False) follow the first marker in the raw text.

  • absent – When True, the expectation is inverted: it passes only when NO event of this type arrives before the within_ms budget expires, and fails as soon as one does. Matches on event type only; text_contains, text_excludes, eval, calls and the marker checks are not allowed alongside it.

within_ms: int | None = None
text_contains: str | None = None
text_excludes: str | None = None
calls: list[EvalFunctionCall] | None = None
eval: str | None = None
marker: str | None = None
marker_first: bool | None = None
markers: int | None = None
text_after: bool | None = None
absent: bool = False
property aggregates: bool

Whether the check accumulates text across events rather than matching one.

A reply with a content check accumulates the bot’s segments; a user_transcription with text_contains accumulates an STT’s pieces.

class pipecat.evals.script.EvalSendAfter(event: str | None, delay_ms: int)[source]

Bases: object

Scheduling for when a turn’s input (user or dtmf) is sent.

When set on a EvalScriptTurn, the harness waits for event to have been seen (either earlier in the run or arriving now), then waits an additional delay_ms before sending the turn’s input. Used for barge-in tests: send_after: {event: llm_started, delay_ms: 500} means “interrupt 500ms after the bot started responding.”

event is optional: a bare send_after: {delay_ms: 500} is a pure time delay with no event anchor (500ms after the previous turn’s send). Handy for pacing keypresses across dtmf turns, where there is no per-key event to anchor on.

Parameters:
  • event – Name of the event to schedule from, or None for a pure delay_ms time delay with no event anchor.

  • delay_ms – Additional delay in milliseconds after the event was received (or, when event is None, after the previous turn’s send).

class pipecat.evals.script.EvalScriptTurn(user: str | None, expect: list[EvalExpectation] = <factory>, audio: str | None = None, dtmf: str | None = None, send_after: EvalSendAfter | None = None, image: str | None = None)[source]

Bases: object

One turn in a scenario.

A turn drives the bot one of three ways: the harness sends a user utterance (the person speaks), it sends a dtmf keypress sequence (the person presses keys), or it is observation-only (neither field — useful for bot-first scenarios like opening greetings). user and dtmf are mutually exclusive: a turn is one or the other.

A user turn is spoken by the harness’s TTS unless the turn names an audio file to play instead.

Parameters:
  • user – Optional text the harness sends as the user’s turn — an RTVI send-text in text modality, or synthesized speech (raw-audio) in audio modality. If absent, the turn just waits for and asserts on expected events.

  • audio – Optional path to an audio file (resolved relative to the scenario file) played as this turn instead of synthesizing user. Any format soundfile reads works (WAV, MP3, FLAC, OGG, …); multi-channel audio is downmixed to mono. user is required alongside it and is what the recording says — the judge reads it as the turn’s input and text_contains matches against it, neither of which can be recovered from the audio. Requires user.modality: audio, and is mutually exclusive with dtmf.

  • dtmf – Optional DTMF keypad sequence the harness sends, one InputDTMFFrame per character (e.g. "123#"). Each character must be a valid KeypadEntry (0-9, *, #). Mutually exclusive with user. The keys are injected the same way regardless of the scenario’s user/judge modality; a bot with a DTMFAggregator turns them into a transcription it reacts to. Quote the value in YAML (dtmf: "123#") — an unquoted # starts a comment.

  • expect – Expected events, in the order they should arrive. Optional — omit it for a pure pacing/observation turn (e.g. a dtmf turn that only presses keys, with the assertion on a later turn).

  • send_after – Optional schedule for when the turn’s input should fire. Only meaningful when user or dtmf is set.

  • image – Optional path to an image to register for this turn (resolved relative to the scenario file). When a function-calling-video bot requests a user image during the turn, the eval transport serves this one. Stays registered until a later turn provides a different image.

audio: str | None = None
dtmf: str | None = None
send_after: EvalSendAfter | None = None
image: str | None = None
class pipecat.evals.script.EvalTurn(**kwargs)[source]

Bases: EvalScriptTurn

Deprecated alias for EvalScriptTurn.

Deprecated since version 1.9.0: Use EvalScriptTurn instead. Will be removed in 2.0.0.

class pipecat.evals.script.EvalScriptScenario(name: str, turns: list[~pipecat.evals.script.EvalScriptTurn], context: list[dict] = <factory>, judge: dict = <factory>, bot_audio: bool = False, transcriber: dict | None = None, user_audio: bool = False, user_speech: dict | None = None, trigger_disconnect: bool = False, stop_on_failure: bool = True, source_path: ~pathlib.Path | None = None)[source]

Bases: object

A parsed scenario file.

Parameters:
  • name – The eval name (from name:).

  • turns – Ordered list of turns.

  • context – LLM messages the bot’s context should start from for this eval. When non-empty, the harness sends them as an eval-context client message right after the bot-ready handshake (the eval serializer turns it into an LLMMessagesUpdateFrame, which replaces the context); bots without an LLM context aggregator ignore the frame. Omitted or empty (the default): the harness sends nothing and the bot keeps the context it set up itself.

  • judge – Judge LLM configuration dict with keys service, model, optional endpoint, and an optional extra mapping forwarded to the model as top-level request parameters. Defaults to {"service": "ollama", "model": "gemma4:12b", "extra": {"reasoning_effort": "none"}}.

  • bot_audio – Whether the bot produces speech, derived from judge.modality. False (text, the default): the bot skips TTS — the harness configures skip-TTS at connect, so even an on-connect greeting is silent. True (audio): the bot speaks, and the judge evaluates the transcription of its actual audio.

  • transcriber – Parsed from the judge.transcription: block; the STT config (service defaults to moonshine, plus model and an optional language code) used to transcribe the bot’s audio for the response event (None in text modality). Set language when the bot speaks a non-English language so the STT doesn’t default to English.

  • user_audio – Whether the user’s turns reach the bot as speech, derived from user.modality. False (text, the default): each turn is sent as an RTVI send-text. True (audio): the harness streams RTVI raw-audio, exercising the bot’s STT for real.

  • user_speech – Parsed from the user.speech: block; the TTS config the harness synthesizes user turns with (None in text modality). Mapping with service, voice, and optional model / language / speed / sample_rate / api_key. Set language (a code like zh) to synthesize non-English user turns.

  • trigger_disconnect – Whether the harness fires the bot’s on_client_disconnected handler when this scenario’s connection ends. Bots often cancel their pipeline there, so this is False by default to avoid that between scenarios; set True to exercise the bot’s disconnect path. Independent of --stop-bot, which tears the bot down via eval-cancel regardless of the handler.

  • stop_on_failure – Whether the first failed turn ends the scenario (default True). A failed turn leaves the conversation in an unknown state, so continuing usually costs one timeout per remaining turn. Set False for a scenario whose turns are scored independently, where the turns after a failure are still worth driving; each turn’s outcome is reported in turns. This governs turn-to-turn progression only: within a turn, an expectation that times out still ends that turn’s matching, because a turn’s expectations share one deadline anchored at the send.

  • source_path – Path the scenario was loaded from, for error messages.

bot_audio: bool = False
transcriber: dict | None = None
user_audio: bool = False
user_speech: dict | None = None
trigger_disconnect: bool = False
stop_on_failure: bool = True
source_path: Path | None = None
classmethod load(path: str | Path) → EvalScriptScenario[source]

Parse a YAML file holding a scenario’s own keys at its top level.

Deprecated since version 1.11.0: Use load() instead. Will be removed in 2.0.0.

Parameters:

path – Path to a YAML file with the scenario schema.

Returns:

The parsed scenario.

wants_response() → bool[source]

Whether any expectation asserts on the transcription of the bot’s audio.

required_report_level() → str | None[source]

The function-call report level the scenario’s assertions need: full for args, name for names, else None.

A call with an eval: needs full too: the judge is asked about the call’s arguments.

needs_vad_events() → bool[source]

Whether the scenario uses the raw VAD speaking events, which the bot emits only on request.

needs_bot_images() → bool[source]

Whether the scenario asserts on the images the bot outputs, which it reports only on request.

needs_marker_events() → bool[source]

Whether the scenario asserts on the LLM’s markers, which the bot emits only on request.

class pipecat.evals.script.EvalScenario(**kwargs)[source]

Bases: EvalScriptScenario

Deprecated alias for EvalScriptScenario.

Deprecated since version 1.9.0: Use EvalScriptScenario instead. Will be removed in 2.0.0.