aggregated_frame_sequencer
Ordered sequencer for AggregatedTextFrame slots through TTS processing.
- class pipecat.utils.context.aggregated_frame_sequencer.AggregatedFrameSequencer(name: str = 'AggregatedFrameSequencer', streaming: bool = False)[source]
Bases:
objectSequences AggregatedTextFrame slots to preserve TTS context ordering.
Manages an ordered queue of spoken and skipped TTS slots. Spoken slots are tracked via a
WordCompletionTracker; skipped slots (e.g. code blocks excluded from TTS synthesis) wait in-place until all preceding spoken slots are complete, then are flushed downstream withappend_to_context=True.Most methods are synchronous and return lists of frames the caller should push downstream, making the sequencer easily testable. The exceptions are
register_spoken(),register_skipped(), andfinalize(), which are async because — when the sequencer is built withstreaming=True— they drive an async_ParallelSentenceAggregatorto group streamed tokens into sentences.Contexts can be in flight concurrently (e.g. two back-to-back
TTSSpeakFrameutterances on a websocket TTS service whoserun_ttsreturns before synthesis finishes). State is organised in three tiers so concurrent contexts never bleed into each other:_slots: the single global ordered timeline across all contexts. Ordering is the whole point of the sequencer, so this stays one list._context_append_to_context: a per-context flag whose presence also marks the context as live — an entry existing is the “process this context’s words” signal inprocess_word()’s stale gate. Created at slot registration, removed when the context is fully done (force_complete())._streaming_contexts: transient per-context pending-sentence state (_StreamingContext), present only while a sentence is being accumulated from tokens. Released at end of text input (finalize()).
Example:
sequencer = AggregatedFrameSequencer() await sequencer.register_spoken(frame, ctx_id, tts_text, append_to_context=True) for f in sequencer.process_word("hello", pts=1000, context_id=ctx_id): await self.push_frame(f)
- __init__(name: str = 'AggregatedFrameSequencer', streaming: bool = False)[source]
Initialize the sequencer.
- Parameters:
name – Label used in log messages (typically the owning TTS service name).
streaming –
True when tokens are dispatched to the TTS individually (
TextAggregationMode.TOKEN). Eachregister_spoken()call then represents one token rather than a complete unit, so tokens are fed to a_ParallelSentenceAggregatorand only turned into a real slot once a sentence boundary is detected (or forced viaregister_skipped()/finalize()). Fixed for the life of the sequencer — a TTS service’s aggregation mode never changes at runtime.Requires the owning TTS service to reuse one context ID for the whole turn (
reuse_context_id_within_turn=True, the default): a promoted sentence built from several tokens is registered under a single context ID, and word-timestamp events for all of its tokens must arrive tagged with that same ID. A per-token context ID would leave every token but the last unmatched, and its words dropped as stale.
- async register_spoken(frame: AggregatedTextFrame, context_id: str, tts_text: str, append_to_context: bool, build_tracker: bool = True, includes_inter_frame_spaces: bool = False) list[Frame][source]
Register a spoken AggregatedTextFrame slot.
Called from _push_tts_frames for every frame sent to the TTS service (one call per token when
streaming=True, one call per complete unit otherwise). Builds theWordCompletionTrackerinternally whenbuild_trackeris set — callers never construct one themselves. A registered slot is marked complete either viaprocess_word()(word-timestamp services) orcomplete_spoken_slot()(push_text_frames=True services).When the sequencer is non-streaming, this registers a slot immediately. When streaming — whether or not
build_trackeris set — the call instead feeds this token to the_ParallelSentenceAggregatorand only registers a real slot once a sentence boundary is confirmed there; a push_text_frames=True (no-tracker) slot completes immediately on promotion (see_promote()) since there is no word-timestamp signal to complete it progressively.- Parameters:
frame – The AggregatedTextFrame being spoken (one token when streaming).
context_id – The TTS context ID assigned to this frame.
tts_text – The text actually sent to the TTS for this call (may differ from
frame.textafter filters/transforms).append_to_context – Whether word frames built for this context should carry append_to_context=True.
build_tracker – Whether to track word completion at all. False for push_text_frames=True services, which complete via complete_spoken_slot instead of word-timestamp matching.
includes_inter_frame_spaces – When True, every TTSTextFrame emitted for this slot carries
includes_inter_frame_spaces=Trueso downstream consumers do not inject extra spaces between consecutive frames. Not used on the streaming path — there, CJK spacing is driven solely byprocess_word()’s per-call flag.
- Returns:
Frames unblocked by this call (a promoted sentence and/or buffered words replayed once a pending sentence promotes). Always empty for the non-streaming case.
- async register_skipped(frame: AggregatedTextFrame, context_id: str, transport_destination: str | None) list[Frame][source]
Register a skipped AggregatedTextFrame and attempt an immediate flush.
Any sentence still pending for this context is finalized first, so a real spoken slot exists immediately before the skipped slot in the queue —
flush()’s “stop at first incomplete spoken slot” logic then blocks this skipped frame correctly until that sentence is actually spoken.The frame is appended as a skipped slot. If no incomplete spoken slot precedes it, the frame is returned right away; otherwise it waits until a later
flush()unblocks it.- Parameters:
frame – The skipped AggregatedTextFrame (e.g. a code block).
context_id – The context ID assigned in _push_tts_frames.
transport_destination – Transport routing value to attach at flush time.
- Returns:
any sentence promoted by the initial
finalize()(streaming mode), followed by this skipped frame once it is unblocked. The skipped frame itself is absent while a preceding spoken slot is still incomplete — the promoted-sentence frame can still be returned in that case, so the list is not necessarily empty when blocked.- Return type:
Frames to push downstream
- async finalize(context_id: str | None) list[Frame][source]
Force-promote a context’s still-pending sentence into a real slot.
Called at end of text input for a context (no more tokens are coming), to handle a response that ends with no terminal punctuation. The context’s pending-sentence state is released here; its
_context_append_to_contextentry is deliberately kept, since word-timestamp events for the just-promoted slot still arrive later (during audio playback) and must be recognised as live. A no-op when nothing is pending for the context (or the sequencer is not streaming).- Parameters:
context_id – The context whose pending sentence should be finalized.
- Returns:
Frames unblocked by finalizing (e.g. buffered words that can now be replayed against the newly-registered slot).
- process_word(word: str, pts: int, context_id: str | None, includes_inter_frame_spaces: bool = False) list[Frame][source]
Process one word-timestamp event and return frames to push downstream.
Locates the active slot (the first incomplete spoken slot with a tracker for this
context_id), advances it by the incoming word, and builds aTTSTextFrame. Handles:Words from a context that was never registered, was wiped by
clear()on interruption, or has already completed (force_complete()removes its entry): dropped as stale (returns an empty list).Normal words that fit entirely within the active slot.
Overflow words straddling two slot boundaries.
Force-complete when the TTS drops an event (word belongs to the next slot).
Passthrough for words not recognised by any slot (buffered instead, when streaming, since the slot they belong to may simply not be promoted yet).
Flushes any skipped slots unblocked by slot completion.
- Parameters:
word – A word token from the TTS service word-timestamp stream.
pts – Presentation timestamp (nanoseconds) to assign to the frame.
context_id – TTS context ID from the word-timestamp event.
includes_inter_frame_spaces – Stamped onto the emitted TTSTextFrame so downstream consumers know not to inject extra spaces between frames.
- Returns:
Ordered list of frames (TTSTextFrame and/or AggregatedTextFrame) to push.
- complete_spoken_slot() list[Frame][source]
Mark the first pending spoken slot complete and flush unblocked skipped frames.
Used by push_text_frames=True services: non-streaming callers invoke this directly after appending the unit’s TTSTextFrame to the audio context;
_promote()calls it internally for a streaming (no-tracker) sentence, immediately on promotion. Only ever targets a build_tracker=False slot (a word-timestamp slot completes viaprocess_word()instead), so “first pending” is always the caller’s own slot even with multiple contexts in flight — such slots complete synchronously, one at a time, in the same call chain that registers them.- Returns:
AggregatedTextFrame(s) that are now unblocked and should be pushed.
- flush(last_word_pts: int | None = None) list[Frame][source]
Walk the slot queue and return all skipped frames that are now unblocked.
Removes complete spoken slots from the head of the queue, then emits (and removes) skipped slots whose preceding spoken slots are all done. Stops at the first incomplete spoken slot.
- Parameters:
last_word_pts – When provided, skipped frames receive this PTS so they appear immediately after the last spoken word in the timeline.
- Returns:
AggregatedTextFrame(s) ready to be pushed downstream.
- force_complete(context_id: str, last_word_pts: int) list[Frame][source]
Force-complete a context’s incomplete spoken slots and flush skipped frames.
Called at the end of an audio context to handle TTS providers that silently drop word-timestamp events. Emits a TTSTextFrame for any remaining unspoken text in each incomplete slot belonging to this context, marks it complete, then flushes all now-unblocked skipped frames. Slots for other contexts still in flight are left untouched so their own word events (or their own force_complete) finish them.
Audio contexts are drained one at a time in registration order (a context’s whole
_handle_audio_contextruns, ending with its own force_complete, before the next starts), so an earlier context is never still open here — every earlier context has already been force-completed by the time this one ends.The context is fully done once this returns, so its live-context and pending state are forgotten — any word that arrives afterwards is dropped as stale.
- Parameters:
context_id – The audio context that has ended.
last_word_pts – PTS of the last received word frame, used as the PTS for force-completed frames and forwarded to
flush().
- Returns:
Combined list of TTSTextFrames (for incomplete spoken slots) and AggregatedTextFrames (skipped slots now unblocked), in emission order.