tts

The user’s voice: a caching TTS in the harness pipeline.

In audio mode the harness synthesizes each user turn and streams it to the bot. CachingTTSService wraps a real TTS service and caches its audio on disk, keyed by service, voice, model, language, speed, and text, so a scripted utterance is synthesized once and reused across runs and bots.

Only local (Kokoro) and HTTP (Cartesia) services fit, since the wrapper reads the inner service’s run_tts generator directly.

pipecat.evals.tts.tts_sample_rate(voice_cfg: dict) → int[source]

The sample rate a user_audio block asks for (default 16 kHz).

pipecat.evals.tts.tts_cache_key(voice_cfg: dict) → str[source]

A stable identity for a user_audio config: service, voice, model, language, and speed, not the sample rate.

class pipecat.evals.tts.CachingTTSService(inner: TTSService, *, cache_key: str, cache_dir: str | Path | None = None, use_cache: bool = True, **kwargs)[source]

Bases: TTSService

A pipeline TTS that wraps a real service and caches its audio on disk.

On a TTSSpeakFrame it emits the cached audio when present, otherwise it runs the wrapped service, caches the audio, and forwards it. The inner service’s lifecycle is forwarded so it works inside the pipeline. Only local and HTTP inner services are supported.

__init__(inner: TTSService, *, cache_key: str, cache_dir: str | Path | None = None, use_cache: bool = True, **kwargs)[source]

Initialize the caching TTS.

Parameters:
  • inner – The local/HTTP TTSService that does the synthesis.

  • cache_key – Stable identity for the inner config (see tts_cache_key()); combined with the text to key the cache.

  • cache_dir – Where to store cached audio. Defaults to <user-cache-dir>/pipecat/evals/tts (or $PIPECAT_EVALS_CACHE_DIR).

  • use_cache – When False, ignore cached audio and don’t write new files.

  • **kwargs – Additional arguments passed to TTSService.

Raises:

ValueError – If inner is a websocket-streaming TTS service (its run_tts doesn’t yield audio); use a local or HTTP service.

async setup(setup)[source]

Set up this service and forward setup to the inner service.

async start(frame: StartFrame)[source]

Start this service and the inner one, the inner with metrics off since it has no downstream.

async stop(frame)[source]

Stop the inner service, then this one.

async cancel(frame)[source]

Cancel the inner service, then this one.

async cleanup()[source]

Clean up the inner service, then this one.

async run_tts(text: str, context_id: str) → AsyncGenerator[Frame | None, None][source]

Emit the audio for text from cache, or by driving the inner service.

Parameters:
  • text – The utterance to synthesize.

  • context_id – TTS context id (forwarded to the inner service).

Yields:

TTSStartedFrame / TTSAudioRawFrame / TTSStoppedFrame (cache hit), or the inner service’s own frames (cache miss).