word_completion_tracker

Per-frame bookkeeping for the words a TTS provider reports speaking.

class pipecat.utils.context.word_completion_tracker.WordCompletionTracker(tts_text: str, llm_text: str | None = None, user_facing_text: str | None = None)[source]

Bases: object

Follows one AggregatedTextFrame from dispatch until it is fully spoken.

A TTS provider reports the words it speaks one event at a time. This class consumes those events for a single frame and answers, for each one:

Three texts describe the same frame, and each answer above is phrased in one of them:

Text

Example

Answers about

tts_text

<spell>4111 1111</spell>

what was spoken

user_facing_text

4111 1111

what a UI displays

llm_text

<card>4111 1111</card>

what the context stores

Keeping a position in all three is the job of TextSegmentMap, which this class owns one of and defers to for every question of where. What the tracker adds is the handful of decisions a position alone cannot express:

  • Providers drop events. A word the provider never reported has no event coming, so the next word that does arrive is matched a little further on and the text passed over is emitted with it. When nothing matches at all, waiting would stall this frame and everything queued behind it, so the frame is force-completed instead: the unspoken remainder is emitted so the context still gets it, and the stray word is handed back for the next frame to try.

  • Some text is never spoken. A closing </card>, or a tag sitting between the last word and its punctuation, never arrives as its own event. Whatever is left once everything speakable is done belongs to this frame, so the word that finishes it claims the rest.

  • A word can belong to two frames. A provider may merge across the boundary ("1111And"). The part that fits stays here; the rest is exposed as overflow for the caller to feed to the next frame.

Example:

tracker = WordCompletionTracker("Hello, world!")
tracker.add_word_and_check_complete("Hello")   # False
tracker.add_word_and_check_complete("world")   # True -- nothing left to speak
__init__(tts_text: str, llm_text: str | None = None, user_facing_text: str | None = None)[source]

Initialize the tracker with the frame’s three texts.

Only tts_text is required; the other two default to it, which is exactly right for a frame nothing rewrote.

Parameters:
  • tts_text – What was sent to the TTS, and so what the incoming words are matched against. May carry synthesis tags (<spell>...).

  • llm_text – What the LLM wrote, with any delimiters an aggregator split off (<card>4111 1111</card>). Supply it to have each word attributed back to it via get_llm_consumed(), which is what keeps those delimiters in the conversation context.

  • user_facing_text – What a client displays – no tags, no rewrites. Defaults to tts_text with markup stripped.

add_word_and_check_complete(word: str) → bool[source]

Record one word the TTS provider reported speaking.

Three things can happen, in this order:

  1. The frame is already finished – the word is ignored.

  2. The word does not match what is left to speak, so the provider must have dropped an event: the frame is force-completed (see _force_complete()) and this word is handed back as overflow.

  3. Otherwise the word advances the frame. Afterwards the get_* accessors describe it: this frame’s share of the word, the LLM text it stands for, and how much of the frame is now spoken.

Parameters:

word – One token from the provider’s word-timestamp stream. It may be a plain word, a word carrying its own spacing or punctuation, or a fragment of a still-open SSML tag – matching is textual, so none of those need special handling from the caller. Services that report spaces and punctuation as separate tokens (e.g. Inworld) must merge them into the preceding word first, via merge_punct_tokens.

Returns:

True once nothing is left for this frame to speak.

take_remaining_as_spoken() → None[source]

Move the cursors to the end, for a caller that emits the remainder itself.

The accumulated and remaining views then agree with the text that went out.

word_belongs_here(word: str) → bool[source]

Return True if word plausibly continues what this frame has left to say.

A False answer means the provider dropped an event. Callers ask this before adding a word so they can offer it to the next frame instead; adding it anyway force-completes this one.

suppress_in_context() → bool[source]

True when the last word must not be written to the conversation context.

Two kinds of word answer to this, both of which the context already has covered, or will have:

  • One step inside a rewritten span. "$42.50" is spoken as five words, none of which the transcript should contain; the word that finishes the span carries "$42.50" for all of them.

  • A word with no span of its own. A provider reporting "," on its own, after "Yeah" took the comma into its span, has nothing left to record. The context falls back to the spoken text when a word carries no span, which would store the mark a second time.

The word is still emitted either way: the provider spoke it, and a consumer reading the word stream should see it. Without an llm_text there are no spans at all, so nothing is suppressed and every word is recorded from its spoken text as usual.

get_word_for_frame() → str | None[source]

Return this frame’s share of the last word – the text to emit for it.

Usually the whole word. A word straddling the boundary gives up its tail ("1111" out of "1111And"), and a word that repeats the previous word’s punctuation gives up that mark. A word matched past an event the provider never sent brings along the text it passed over, and after a force-complete this is the frame’s unspoken remainder instead – either way, nothing is missing from the turn.

get_overflow_word() → str | None[source]

Return the part of the last word that belongs to the next frame.

Feed it to that frame’s tracker as if the provider had sent it there. Casing and punctuation are untouched so it still reads as a real word. None when the word fit entirely within this frame.

get_llm_consumed() → str | None[source]

Return the LLM’s own text that the last word stands for.

This is what the conversation context records, so it keeps the tags and spellings the LLM wrote ("<card>4111") rather than what the provider reported speaking ("4111").

None when there is nothing to attribute: no llm_text was given, the word is mid-rewrite, or llm_text is already exhausted (a trailing emoji it never carried).

get_accumulated_user_facing_text() → str[source]

Return the part of the frame spoken so far, as the user sees it.

With get_remaining_user_facing_text() this splits the frame’s text in two, which is what lets a client highlight speech as it happens.

get_remaining_user_facing_text(strip: bool = True) → str[source]

Return the part of the frame not yet spoken, as the user sees it.

Parameters:

strip – Whether to trim surrounding whitespace. Pass False to keep the leading space, so accumulated + remaining reproduces the frame’s text exactly – callers that index into that text rely on it.

get_accumulated_tts_text() → str[source]

Return everything spoken so far, as it was sent to the TTS.

The whole frame up to the cursor, where get_word_for_frame() describes only the most recent word.

get_accumulated_llm_text() → str | None[source]

Return everything spoken so far, as the LLM wrote it.

The whole frame up to the cursor, where get_llm_consumed() describes only the most recent word. None without an llm_text.

get_remaining_tts_text(strip: bool = True) → str[source]

Return what this frame still has left to speak.

Callers ending a frame early emit this, so text the provider never reported still reaches the conversation context.

Parameters:

strip – Whether to trim surrounding whitespace. Pass False to keep the leading space, so accumulated + remaining reproduces tts_text exactly.

get_remaining_llm_text() → str | None[source]

Return what this frame still has left to speak, as the LLM wrote it.

The companion to get_remaining_tts_text() when ending a frame early: that supplies the text to emit, this the text to record. None without an llm_text, or when nothing is left.

property is_complete: bool

True when this frame has nothing left to speak.

Alphanumeric content is what counts, so a frame whose remainder is only punctuation or a closing tag is already finished – no word will ever arrive for those.

reset()[source]

Rewind to the start of the frame, keeping the three texts.