text_segment_map
Keeps a position in three versions of the same text as a TTS speaks it.
- class pipecat.utils.context.text_segment_map.TextSegment(original: str, tts: str, original_start: int, original_end: int)[source]
Bases:
objectA piece of the utterance, paired with what the TTS was given in its place.
The map is a list of these, laid end to end over the whole utterance. In most of them the two sides are identical; the interesting ones are where a transform, a filter or a tag made them differ (see
is_transformed).- Parameters:
original – The piece as a client displays it.
tts – The same piece as it was sent to the TTS. Identical to original unless something rewrote it, and empty if it was dropped entirely.
original_start – Where the piece starts in the full original text.
original_end – Where it ends. Cursors jump straight here once a rewritten piece is finished, since no position inside one means anything.
- property is_transformed: bool
True when the two sides cannot be followed together, character by character.
A segment like this is all or nothing. The cursors into the other texts wait at its start until every spoken word of it has arrived, then jump straight to its end, because no position inside it means anything.
Any one of these makes it true:
the letters and digits differ, as in
"$42.50"against"forty two dollars";the two sides have different numbers of words;
the TTS side has tags in it, even when the words match. The cursor through the spoken text has tag characters to cross that the other texts do not have.
Only the shape of a tag matters, never its name:
<phoneme ...>Siobhan</phoneme>counts because of the tags around the word.
- property tts_alnum_count: int
How many letters and digits the spoken side of this segment has.
This is what the provider’s words spend as they arrive. On a rewritten segment it has nothing to do with
original_alnum_count: “forty two dollars” against “$42.50”.
- class pipecat.utils.context.text_segment_map.TextSegmentMap(tts_text: str, original_text: str, llm_text: str | None = None)[source]
Bases:
objectAnswers “where are we?” in three versions of one utterance, word by word.
A TTS provider reports the words it speaks. Each report has to be turned into a position – but into a position in three different strings, because the same utterance exists in three forms at once:
tts_text– what was actually spoken, tags and all:"Your balance is forty two dollars"original_text– what a client displays:"Your balance is $42.50"llm_text– what the LLM wrote, so what the transcript should keep:"Your balance is <b>$42.50</b>". Defaults tooriginal_text.
For a frame nothing rewrote, all three are the same string and every position is the same.
The hard part is that a spoken word need not appear in the other two. The provider says
"dollars"; nothing in"$42.50"matches it. So the map is built once, by diffingtts_textagainstoriginal_textinto alignedTextSegmentpieces – each either survived unchanged or was rewritten whole.llm_textis never compared against the others, and does not need to be. It holds the same letters and digits asoriginal_text, in the same order, and differs only in what is wrapped around them – tags, delimiters, punctuation. So counting letters and digits is enough to keep it in step, and its cursor moves by that count.From then on one real cursor moves:
raw_pos, how far intotts_textthe provider has got.user_facing_posand the two LLM cursors follow it:Through an unchanged segment they keep pace, word for word.
Through a rewritten one they wait. There is no honest position halfway through
"$42.50"while"forty two dollars"is being spoken, so they hold and then jump to the end of the span in one step when the last of its words lands.
The LLM side carries two of them because one position cannot answer both questions asked of it.
llm_spoken_posis what the provider has reported;llm_posis what has been attributed to a word, and runs ahead of it over a mark stuck to a word the provider has already named. See_advance_llm_cursors().Callers ask two things.
word_belongs_current_segment()– does this token plausibly continue what is left to speak? – andadvance_word(), which consumes it. Both tolerate the ways providers mangle tokens (added punctuation, changed case or diacritics, a fragment of a half-open SSML tag) without the caller knowing anything about it;_classify_hop()holds that logic.A word need not sit at the cursor to be placed. When the provider garbles or drops an event, the next good word still matches a few words on, and the text stepped over – which no event will ever name – is consumed with it. So the two questions above stay simple: a token either has a home somewhere in what is left to speak, or it belongs to another utterance entirely.
Example:
# "$42.50" was sent to the TTS as "forty two dollars and fifty cents" smap = TextSegmentMap( "Your balance is forty two dollars and fifty cents", "Your balance is $42.50", ) for word in ["Your", "balance", "is"]: smap.advance_word(word) # unchanged: every cursor keeps pace for word in ["forty", "two", "dollars", "and", "fifty"]: smap.advance_word(word) # rewritten: the other two cursors wait smap.advance_word("cents") # the span is done, so they jump to its end assert smap.last_completed_segment.original == "$42.50" assert not smap.in_transformed_segment
- LOOKAHEAD_WORDS = 3
How many words a word may be placed past – see
_lookahead_hop().
- __init__(tts_text: str, original_text: str, llm_text: str | None = None)[source]
Line the three texts up against each other.
The comparison happens once, here. Everything after this only moves cursors.
- Parameters:
tts_text – What was sent to the TTS, and so what incoming words are matched against. May carry synthesis tags and rewritten values.
original_text – The same content as a client displays it, before any rewriting. Diffed against tts_text to build the segments.
llm_text – The same content as the LLM wrote it, which may add delimiters the other two never see. Rides its own cursor rather than being diffed. Defaults to original_text.
- advance_word(word: str) None[source]
Take one spoken word and move every cursor to where it ends.
Afterwards
last_completed_segment,last_overflowandlast_leading_duplicatedescribe what this particular word did; each is cleared at the start of the next call.- Parameters:
word – One token from the provider’s word-timestamp stream. It may be a plain word, a word carrying its own spacing or punctuation, or a fragment of a half-open tag – matching is textual, so the caller does not have to know which.
- word_belongs_current_segment(word: str) bool[source]
Return True if word could be the next thing spoken here.
advance_word()without the moving, so a caller can check first. A False answer means the provider skipped ahead, and the word should go to the next frame instead.A word with no letters or digits gets a second chance from
_symbol_belongs_here(), since there is nothing in it to match on.
- property llm_pos: int
How far into the LLM’s text has been attributed to a word so far.
Ahead of
llm_spoken_poswhenever a word swept up a mark that no event has reported yet; the text between the two is attributed but not yet spoken.
- property llm_spoken_pos: int
How far into the LLM’s text the provider has actually reported speaking.
What a caller showing progress, or ending a frame early, should read: text past this point has had no word event, so it still belongs to what is left to say even when
llm_poshas already credited it to a word.
- property last_overflow: str | None
The end of the last word passed to
advance_word(), if it did not fit.Nonemost of the time. It is set only when that word ran past the end oftts_text, with no segment left to take the rest, which means the leftover belongs to the next frame. It is always the tail of the word that was passed in, so the part that did fit isword[: len(word) - len(last_overflow)].
- property last_leading_duplicate: int
How much of the last word’s start was punctuation already spoken.
The opposite end of the word from
last_overflow: that one is about a tail running past this text, this one about a head repeating punctuation the previous word already took. Cut both off to get the part of the word that belongs to this frame:word[last_leading_duplicate : len(word) - len(last_overflow or "")]
- property is_complete: bool
True once every letter and digit in the text has been spoken.
This is not the same as the cursor reaching the end. If all that is left is punctuation or tags, the text counts as finished even though those characters have not been walked over, because no word event is coming for them.
There is one exception. Punctuation separated from its word by a space, as French writes
"Comment ça va ?", does arrive as its own word event, so the text stays unfinished until it does (see_pending_separated_punctuation()). Punctuation stuck to the word itself, as in"you?", was already taken with the word.
- property last_completed_segment: TextSegment | None
The segment finished by the last
advance_word()call, if any.