text_segment_map

Keeps a position in three versions of the same text as a TTS speaks it.

class pipecat.utils.context.text_segment_map.TextSegment(original: str, tts: str, original_start: int, original_end: int)[source]

Bases: object

A piece of the utterance, paired with what the TTS was given in its place.

The map is a list of these, laid end to end over the whole utterance. In most of them the two sides are identical; the interesting ones are where a transform, a filter or a tag made them differ (see is_transformed).

Parameters:
  • original – The piece as a client displays it.

  • tts – The same piece as it was sent to the TTS. Identical to original unless something rewrote it, and empty if it was dropped entirely.

  • original_start – Where the piece starts in the full original text.

  • original_end – Where it ends. Cursors jump straight here once a rewritten piece is finished, since no position inside one means anything.

property is_transformed: bool

True when the two sides cannot be followed together, character by character.

A segment like this is all or nothing. The cursors into the other texts wait at its start until every spoken word of it has arrived, then jump straight to its end, because no position inside it means anything.

Any one of these makes it true:

  • the letters and digits differ, as in "$42.50" against "forty two dollars";

  • the two sides have different numbers of words;

  • the TTS side has tags in it, even when the words match. The cursor through the spoken text has tag characters to cross that the other texts do not have.

Only the shape of a tag matters, never its name: <phoneme ...>Siobhan</phoneme> counts because of the tags around the word.

property tts_alnum_count: int

How many letters and digits the spoken side of this segment has.

This is what the provider’s words spend as they arrive. On a rewritten segment it has nothing to do with original_alnum_count: “forty two dollars” against “$42.50”.

property original_alnum_count: int

How many letters and digits the original side of this segment has.

This is what the cursors into original_text and llm_text spend, since those texts hold the original characters, not the spoken ones.

class pipecat.utils.context.text_segment_map.TextSegmentMap(tts_text: str, original_text: str, llm_text: str | None = None)[source]

Bases: object

Answers “where are we?” in three versions of one utterance, word by word.

A TTS provider reports the words it speaks. Each report has to be turned into a position – but into a position in three different strings, because the same utterance exists in three forms at once:

  • tts_text – what was actually spoken, tags and all: "Your balance is forty two dollars"

  • original_text – what a client displays: "Your balance is $42.50"

  • llm_text – what the LLM wrote, so what the transcript should keep: "Your balance is <b>$42.50</b>". Defaults to original_text.

For a frame nothing rewrote, all three are the same string and every position is the same.

The hard part is that a spoken word need not appear in the other two. The provider says "dollars"; nothing in "$42.50" matches it. So the map is built once, by diffing tts_text against original_text into aligned TextSegment pieces – each either survived unchanged or was rewritten whole.

llm_text is never compared against the others, and does not need to be. It holds the same letters and digits as original_text, in the same order, and differs only in what is wrapped around them – tags, delimiters, punctuation. So counting letters and digits is enough to keep it in step, and its cursor moves by that count.

From then on one real cursor moves: raw_pos, how far into tts_text the provider has got. user_facing_pos and the two LLM cursors follow it:

  • Through an unchanged segment they keep pace, word for word.

  • Through a rewritten one they wait. There is no honest position halfway through "$42.50" while "forty two dollars" is being spoken, so they hold and then jump to the end of the span in one step when the last of its words lands.

The LLM side carries two of them because one position cannot answer both questions asked of it. llm_spoken_pos is what the provider has reported; llm_pos is what has been attributed to a word, and runs ahead of it over a mark stuck to a word the provider has already named. See _advance_llm_cursors().

Callers ask two things. word_belongs_current_segment() – does this token plausibly continue what is left to speak? – and advance_word(), which consumes it. Both tolerate the ways providers mangle tokens (added punctuation, changed case or diacritics, a fragment of a half-open SSML tag) without the caller knowing anything about it; _classify_hop() holds that logic.

A word need not sit at the cursor to be placed. When the provider garbles or drops an event, the next good word still matches a few words on, and the text stepped over – which no event will ever name – is consumed with it. So the two questions above stay simple: a token either has a home somewhere in what is left to speak, or it belongs to another utterance entirely.

Example:

# "$42.50" was sent to the TTS as "forty two dollars and fifty cents"
smap = TextSegmentMap(
    "Your balance is forty two dollars and fifty cents",
    "Your balance is $42.50",
)
for word in ["Your", "balance", "is"]:
    smap.advance_word(word)   # unchanged: every cursor keeps pace
for word in ["forty", "two", "dollars", "and", "fifty"]:
    smap.advance_word(word)   # rewritten: the other two cursors wait
smap.advance_word("cents")    # the span is done, so they jump to its end
assert smap.last_completed_segment.original == "$42.50"
assert not smap.in_transformed_segment
LOOKAHEAD_WORDS = 3

How many words a word may be placed past – see _lookahead_hop().

__init__(tts_text: str, original_text: str, llm_text: str | None = None)[source]

Line the three texts up against each other.

The comparison happens once, here. Everything after this only moves cursors.

Parameters:
  • tts_text – What was sent to the TTS, and so what incoming words are matched against. May carry synthesis tags and rewritten values.

  • original_text – The same content as a client displays it, before any rewriting. Diffed against tts_text to build the segments.

  • llm_text – The same content as the LLM wrote it, which may add delimiters the other two never see. Rides its own cursor rather than being diffed. Defaults to original_text.

advance_word(word: str) → None[source]

Take one spoken word and move every cursor to where it ends.

Afterwards last_completed_segment, last_overflow and last_leading_duplicate describe what this particular word did; each is cleared at the start of the next call.

Parameters:

word – One token from the provider’s word-timestamp stream. It may be a plain word, a word carrying its own spacing or punctuation, or a fragment of a half-open tag – matching is textual, so the caller does not have to know which.

word_belongs_current_segment(word: str) → bool[source]

Return True if word could be the next thing spoken here.

advance_word() without the moving, so a caller can check first. A False answer means the provider skipped ahead, and the word should go to the next frame instead.

A word with no letters or digits gets a second chance from _symbol_belongs_here(), since there is nothing in it to match on.

property user_facing_pos: int

How far into the user-facing text the spoken words have reached.

property llm_pos: int

How far into the LLM’s text has been attributed to a word so far.

Ahead of llm_spoken_pos whenever a word swept up a mark that no event has reported yet; the text between the two is attributed but not yet spoken.

property llm_spoken_pos: int

How far into the LLM’s text the provider has actually reported speaking.

What a caller showing progress, or ending a frame early, should read: text past this point has had no word event, so it still belongs to what is left to say even when llm_pos has already credited it to a word.

property raw_pos: int

How far into tts_text the provider has spoken, counted from its start.

property last_overflow: str | None

The end of the last word passed to advance_word(), if it did not fit.

None most of the time. It is set only when that word ran past the end of tts_text, with no segment left to take the rest, which means the leftover belongs to the next frame. It is always the tail of the word that was passed in, so the part that did fit is word[: len(word) - len(last_overflow)].

property last_leading_duplicate: int

How much of the last word’s start was punctuation already spoken.

The opposite end of the word from last_overflow: that one is about a tail running past this text, this one about a head repeating punctuation the previous word already took. Cut both off to get the part of the word that belongs to this frame:

word[last_leading_duplicate : len(word) - len(last_overflow or "")]
property is_complete: bool

True once every letter and digit in the text has been spoken.

This is not the same as the cursor reaching the end. If all that is left is punctuation or tags, the text counts as finished even though those characters have not been walked over, because no word event is coming for them.

There is one exception. Punctuation separated from its word by a space, as French writes "Comment ça va ?", does arrive as its own word event, so the text stays unfinished until it does (see _pending_separated_punctuation()). Punctuation stuck to the word itself, as in "you?", was already taken with the word.

property in_transformed_segment: bool

True when the cursor is partway through a rewritten segment.

property last_completed_segment: TextSegment | None

The segment finished by the last advance_word() call, if any.

reset() → None[source]

Put every cursor back to the start of the text.