ui

UI worker: an LLM worker that observes and drives a client GUI over RTVI.

Composes the RTVI UI wire protocol (client events, accessibility snapshots, server UI commands) with screen_tools, the tool a voice LLM uses to ask a UIWorker about the screen. PipelineWorker connects a UIWorker to the client automatically whenever RTVI is enabled — no decorator or separate component to wire up.

class pipecat.workers.ui.BaseUIWorker(**kwargs)[source]

Bases: BaseWorker

Worker that surfaces its jobs and job groups on the client UI.

Deprecated since version 1.12.0: Use UIWorker instead, which reports its job groups to the client itself. Will be removed in 2.0.0.

Every group this worker dispatches is registered for lifecycle forwarding: a group_started envelope is published at dispatch, worker updates and responses are forwarded as job_update / job_completed envelopes, group_completed is published at group teardown (normal completion, cancellation, or timeout), and the client’s reserved __cancel_job_group event is translated into cancel_job_group for groups dispatched as cancellable.

Instantiable directly (no LLM): register one on the runner as a dispatcher when a pipeline app wants client-visible background work:

ui_jobs = BaseUIWorker("ui-jobs")
job_id = await ui_jobs.request_job_group(
    "wikipedia", "news",
    params=JobGroupParams(
        payload={"query": query},
        label=f"Research: {query}",
    ),
)
async create_job_group_and_request_job(worker_names: list[str], **kwargs) → JobGroup[source]

Dispatch a job group and announce it to the client.

Parameters:
Returns:

The created JobGroup.

async cancel_job_group(job_id: str, *, reason: str | None = None) → None[source]

Cancel a running job group and complete its client card.

Parameters:
  • job_id – The job identifier to cancel.

  • reason – Optional human-readable reason for cancellation.

async on_bus_message(message: BusMessage) → None[source]

Handle the client’s reserved __cancel_job_group event.

Everything else this worker forwards to the client hangs off the job hooks (on_job_update(), on_job_response(), on_job_stream_end(), on_job_completed()), which the base class calls at the right point in a group’s lifecycle.

Parameters:

message – The BusMessage to process.

async on_job_update(message: BusJobUpdateMessage | BusJobUpdateUrgentMessage) → None[source]

Forward a worker’s progress update to the client.

async on_job_response(message: BusJobResponseMessage | BusJobResponseUrgentMessage) → None[source]

Forward a worker’s response to the client as its terminal envelope.

Runs before the group is torn down, so on an error status (with cancel_on_error) the client learns which worker failed before the card closes.

async on_job_stream_end(message: BusJobStreamEndMessage) → None[source]

Forward a worker’s stream end as its terminal envelope.

A worker may finish by ending its stream instead of responding; the client is told it completed, with the final stream data as the response payload.

async on_job_completed(result: JobGroupResponse) → None[source]

Complete the client’s card for a group whose workers all finished.

class pipecat.workers.ui.BusUICommandMessage(command_name: str = '', payload: Any = None, *, source: str, target: str | None = None)[source]

Bases: BusUIDataMessage

A UI command sent from a server-side worker to the client.

Published by UIWorker.send_command(name, payload). The client-facing worker translates it into a command on its client’s wire format; over RTVI that is an RTVIUICommandFrame(command=command_name, payload=payload) pushed through the pipeline.

Parameters:
  • command_name – App-defined command name.

  • payload – App-defined payload (already a plain dict by the time it lands on the bus).

command_name: str = ''
payload: Any = None
class pipecat.workers.ui.BusUIEventMessage(event_name: str = '', payload: Any = None, *, source: str, target: str | None = None)[source]

Bases: BusUIDataMessage

A UI event sent from the client to a server-side worker.

Emitted by the client-facing worker when the client dispatches an event via PipecatClient.sendUIEvent(event, payload). UIWorker subclasses dispatch these to @ui_event(name) handlers.

Parameters:
  • event_name – App-defined event name.

  • payload – App-defined payload. Schemaless by design.

event_name: str = ''
payload: Any = None
class pipecat.workers.ui.BusUIJobCompletedMessage(job_id: str = '', worker_name: str = '', status: str = '', response: Any = None, at: int = 0, *, source: str, target: str | None = None)[source]

Bases: BusUIDataMessage

A worker in a user-facing job group has completed.

Forwarded by a UIWorker when a worker of one of its job groups reaches a terminal state. The client-facing worker forwards it as a ui-job-group envelope with kind = "job_completed".

Parameters:
  • job_id – The shared job-group identifier.

  • worker_name – The worker that produced the response.

  • status – Completion status as a string (JobStatus value).

  • response – The worker’s response payload.

  • at – Epoch milliseconds when the response was received.

job_id: str = ''
worker_name: str = ''
status: str = ''
response: Any = None
at: int = 0
class pipecat.workers.ui.BusUIJobGroupCompletedMessage(job_id: str = '', at: int = 0, *, source: str, target: str | None = None)[source]

Bases: BusUIDataMessage

A user-facing job group has completed.

Published by a UIWorker once every worker in the group has finished, or the group was cancelled. The client-facing worker forwards it as a ui-job-group envelope with kind = "group_completed".

Parameters:
  • job_id – The shared job-group identifier.

  • at – Epoch milliseconds when the group completed.

job_id: str = ''
at: int = 0
class pipecat.workers.ui.BusUIJobGroupStartedMessage(job_id: str = '', workers: list[str] | None = None, label: str | None = None, cancellable: bool = True, at: int = 0, *, source: str, target: str | None = None)[source]

Bases: BusUIDataMessage

A user-facing job group has been dispatched.

Published by a UIWorker as it dispatches the group. The client-facing worker forwards it as a ui-job-group envelope with kind = "group_started".

Parameters:
  • job_id – Shared job-group identifier for the group.

  • workers – Names of the workers the work was dispatched to.

  • label – Optional human-readable label for the group.

  • cancellable – Whether the client may request cancellation.

  • at – Epoch milliseconds when the group started.

job_id: str = ''
workers: list[str] | None = None
label: str | None = None
cancellable: bool = True
at: int = 0
class pipecat.workers.ui.BusUIJobUpdateMessage(job_id: str = '', worker_name: str = '', data: Any = None, at: int = 0, *, source: str, target: str | None = None)[source]

Bases: BusUIDataMessage

Per-worker progress for a user-facing job group.

Forwarded by a UIWorker whenever a worker of one of its job groups emits a BusJobUpdateMessage. The client-facing worker forwards it as a ui-job-group envelope with kind = "job_update".

Parameters:
  • job_id – The shared job-group identifier.

  • worker_name – The worker that produced the update.

  • data – The worker’s update payload, forwarded verbatim.

  • at – Epoch milliseconds when the update was emitted on the bus.

job_id: str = ''
worker_name: str = ''
data: Any = None
at: int = 0
class pipecat.workers.ui.ReplyToolMixin(**kwargs)[source]

Bases: object

Expose a reply tool covering the full standard action set.

Deprecated since version 1.12.0: Use screen_tools() instead: the voice LLM asks the worker through the screen job and says the answer itself. Will be removed in 2.0.0.

Single bundled LLM tool with a required spoken answer plus optional visual and state-changing actions. One tool call per turn, no chaining; the required answer argument is enforced by the API schema so the model cannot omit the terminator.

Compose alongside UIWorker:

class MyUIWorker(ReplyToolMixin, UIWorker):
    ...

Covers pointing apps (scroll_to + highlight), reading apps (scroll_to + select_text), form apps (fills + click), and any blend (e.g. a document review with selection-based deixis AND voice-driven note-taking). The LLM uses whichever fields fit the user’s request per turn; unused fields stay null and don’t affect behavior.

Delivers answer as verbatim TTS (respond_to_job(answer, tts_speak=True)) – the worker speaks the exact phrase. Apps that want a minimal schema (only the fields actually used, or app-specific commands), or that want the requester’s voice LLM to phrase the reply instead, write their own @tool reply on the UIWorker subclass directly. Use the helper methods on UIWorker plus send_command to dispatch the underlying UI commands.

The host class must provide scroll_to, highlight, select_text, click, set_input_value, and respond_to_job (UIWorker does) and must be the target of @tool discovery on the LLM pipeline.

async reply(params: FunctionCallParams, answer: str, scroll_to: str | None = None, highlight: list[str] | None = None, select_text: str | None = None, fills: list[dict] | None = None, click: list[str] | None = None)[source]

Reply to the user. Optionally point at content and act on inputs.

Always called exactly once per turn. answer is required; the action fields are optional and may be combined.

Visual / pointing actions (draw the user’s attention):

  • scroll_to brings an element into view (single ref).

  • highlight flashes elements briefly (list of refs). Best for short emphasis like a button or a fact.

  • select_text puts the page’s text selection on an element (single ref). Best for “this paragraph” / “the section about X” so the user sees exactly what was meant. Persists until the user clicks elsewhere.

State-changing actions (modify form / app state):

  • fills writes values into inputs (list of {"ref", "value"} objects, multi-fill in one turn).

  • click clicks elements (list of refs in order). Use for checkboxes, radios, submit buttons.

Order of dispatch within a turn: scroll_to, then highlight, then select_text, then fills, then click, then speak the answer.

Parameters:
  • params – Framework-provided tool invocation context.

  • answer – The spoken reply in plain language. One short sentence. No markdown, no symbols.

  • scroll_to – Optional snapshot ref. Scrolls the element into view before speaking.

  • highlight – Optional list of snapshot refs. Visually pulses each element.

  • select_text – Optional snapshot ref. Places the page’s text selection on that element.

  • fills – Optional list of {"ref": "eN", "value": "..."} objects. Writes each value into the input at ref.

  • click – Optional list of snapshot refs to click in order.

class pipecat.workers.ui.UISelection(ref: str, text: str)[source]

Bases: NamedTuple

The text the user has selected on the page, and the element it is in.

ref: str

Alias for field number 0

text: str

Alias for field number 1

class pipecat.workers.ui.UIWorker(name: str, *, llm: LLMService[Any], context: LLMContext | None = None, classifier: BaseClassifier | None = None, assistant_params: LLMAssistantAggregatorParams | None = None, inject_events: bool = True, auto_inject_ui_state: bool = True, keep_history: bool = False, prompt_guide: str | None = '## UI context\n\nYour developer context includes two kinds of SDK-managed messages:\n\n- ``<ui_event name="..." >payload</ui_event>``: an event the user just triggered on the client (click, tab switch, navigation, etc.). The payload is JSON for that event.\n- ``<ui_state>...</ui_state>``: an accessibility snapshot of the current screen, injected at the start of every turn. Indented tree in Playwright-MCP style. Each line is ``- role "name" [state] [ref=eN]`` with children nested one level deeper. A line can also carry ``= "value"`` (an element\'s current value, e.g. text already typed into an input) and ``[level=N]`` (heading depth).\n\nState tags include ``[focused]``, ``[selected]``, ``[disabled]``, and ``[offscreen]``. A node tagged ``[offscreen]`` exists on the page but is not currently in the user\'s viewport; only visible (non-offscreen) nodes count for position-based references.\n\nGrids carry a ``[cols=N]`` tag. Their cells are listed in reading order (left-to-right, top-to-bottom); with N columns, cell K sits at row ``ceil(K/N)``, column ``((K-1) mod N) + 1``. Example with ``[cols=8]`` and 16 children: "top right" is cell 8, "bottom left" is cell 9.\n\nResolve position references ("top right", "the first one", "the third new release") against the most recent ``<ui_state>`` tree. Sibling order matches reading order on screen (top-to-bottom, left-to-right within each region).\n\nWhen the user has text selected on the page, the snapshot ends with a ``<selection ref="eN">selected text</selection>`` block inside ``<ui_state>``. Treat the selection as the deictic referent for "this", "that", "what I selected", and similar phrases. The ``ref`` identifies the closest enclosing element that has a ref in the tree; the inner text is the actual selected content (truncated if very long). Text inside ``<input>`` or ``<textarea>`` selections is faithful to ``selectionStart``/``selectionEnd`` on the element.\n\nRefs (``e42``) are stable handles for acting on elements: pass the ``ref`` from the most recent ``<ui_state>`` to any tool that operates on a node. The same element keeps its ref across snapshots while it stays on the page, so you can refer back to it across turns. Always resolve refs against the latest snapshot, and bring an ``[offscreen]`` element into view before acting on it.')[source]

Bases: LLMContextWorker

LLM worker that reads and drives a client GUI over the RTVI UI channel.

A UIWorker connects an LLM to whatever the user is looking at: it sees the screen as accessibility snapshots, reacts to the user’s UI events, and acts on the page by sending commands to the client. It is the delegate side of a voice/UI split – a voice layer (the main pipeline’s LLM, or a separate LLMWorker) handles speech and hands screen-relevant work to this worker.

Capabilities:

  • See the screen. The latest accessibility snapshot is rendered as <ui_state> and auto-injected into the LLM context before each inference. Code reads it through snapshot and the user’s selection.

  • React to UI events, dispatched to @ui_event(name) handlers.

  • Drive the UI with send_command and the scroll_to / highlight / select_text / click / set_input_value helpers.

  • Decide small things with a classifier, not an LLM turn: whether a UI event deserves a comment (should_respond), which element on screen the user means (which_element), whether something is true of the screen (check_screen), which elements match a description (select_elements), and act on an element named in words.

  • Answer a voice LLM’s questions about the screen through the screen job: find an element, check whether something is true, select the elements matching a description, list what is on screen, read the user’s selection, or click, scroll to, highlight, select or fill an element. Every answer is short data and never the page; screen_tools() gives the voice LLM the tool that sends it.

  • Answer as a delegate. The built-in single-flight respond job runs one screen-grounded LLM turn and answers with the reply the LLM writes. A @tool that calls respond_to_job answers instead when it needs to decide how the answer reaches the user.

  • Surface long work. Every job group this worker dispatches is reported to the client as it goes: a card when the group starts, a line per worker’s progress and completion, and the close when the group completes, whether normally, by cancellation or by timeout. The client can cancel a group dispatched as cancellable.

PipelineWorker connects a UIWorker to the client automatically when RTVI is enabled – no extra wiring. A working worker needs only an LLM; subclass it to react to UI events or add jobs and tools, and override render_query to read a non-default job payload.

Example:

class MyUIWorker(UIWorker):
    @ui_event("nav_click")
    async def on_nav(self, message):
        view = message.payload.get("view")
        ...

worker = MyUIWorker("ui", llm=OpenAILLMService(api_key="..."))

Note

With client trackViewport on (the default), off-screen nodes carry [offscreen] in <ui_state>; scroll_to before acting on them.

__init__(name: str, *, llm: LLMService[Any], context: LLMContext | None = None, classifier: BaseClassifier | None = None, assistant_params: LLMAssistantAggregatorParams | None = None, inject_events: bool = True, auto_inject_ui_state: bool = True, keep_history: bool = False, prompt_guide: str | None = '## UI context\n\nYour developer context includes two kinds of SDK-managed messages:\n\n- ``<ui_event name="..." >payload</ui_event>``: an event the user just triggered on the client (click, tab switch, navigation, etc.). The payload is JSON for that event.\n- ``<ui_state>...</ui_state>``: an accessibility snapshot of the current screen, injected at the start of every turn. Indented tree in Playwright-MCP style. Each line is ``- role "name" [state] [ref=eN]`` with children nested one level deeper. A line can also carry ``= "value"`` (an element\'s current value, e.g. text already typed into an input) and ``[level=N]`` (heading depth).\n\nState tags include ``[focused]``, ``[selected]``, ``[disabled]``, and ``[offscreen]``. A node tagged ``[offscreen]`` exists on the page but is not currently in the user\'s viewport; only visible (non-offscreen) nodes count for position-based references.\n\nGrids carry a ``[cols=N]`` tag. Their cells are listed in reading order (left-to-right, top-to-bottom); with N columns, cell K sits at row ``ceil(K/N)``, column ``((K-1) mod N) + 1``. Example with ``[cols=8]`` and 16 children: "top right" is cell 8, "bottom left" is cell 9.\n\nResolve position references ("top right", "the first one", "the third new release") against the most recent ``<ui_state>`` tree. Sibling order matches reading order on screen (top-to-bottom, left-to-right within each region).\n\nWhen the user has text selected on the page, the snapshot ends with a ``<selection ref="eN">selected text</selection>`` block inside ``<ui_state>``. Treat the selection as the deictic referent for "this", "that", "what I selected", and similar phrases. The ``ref`` identifies the closest enclosing element that has a ref in the tree; the inner text is the actual selected content (truncated if very long). Text inside ``<input>`` or ``<textarea>`` selections is faithful to ``selectionStart``/``selectionEnd`` on the element.\n\nRefs (``e42``) are stable handles for acting on elements: pass the ``ref`` from the most recent ``<ui_state>`` to any tool that operates on a node. The same element keeps its ref across snapshots while it stays on the page, so you can refer back to it across turns. Always resolve refs against the latest snapshot, and bring an ``[offscreen]`` element into view before acting on it.')[source]

Initialize the UIWorker.

Parameters:
  • name – Unique name for this worker.

  • llm – The LLM service.

  • context – Optional pre-built LLMContext. Seeded messages are part of the mutable history and are cleared on each keep_history=False reset; put durable instructions in the LLM’s system_instruction instead.

  • classifier – Answers the small questions about the screen (should_respond, which_element). Without one, the llm answers them through an LLMClassifier, which costs an LLM call per question; a JevClassifier answers in about a tenth of a second with a calibrated probability.

  • assistant_params – Optional assistant-aggregator parameters, e.g. to enable context summarization for keep_history=True workers.

  • inject_events – When True (the default), append each UI event to the context as a <ui_event> developer message. Override render_ui_event to change the content, or set False to disable.

  • auto_inject_ui_state – When True (the default), append the latest <ui_state> snapshot to the context before every inference (via the LLM’s on_before_process_frame hook). Set False to inject manually with inject_ui_state().

  • keep_history – When False (the default), the context is cleared at the start of every job, so each turn sees only the current <ui_state> and query – best for the stateless-delegate role. When True, history accumulates across jobs so the LLM can resolve multi-turn references (“the next one”, “the Pro version”), at the cost of more tokens and possible confusion from stale <ui_state> blocks. Use context summarization to prune the history when it gets too large.

  • prompt_guide – Wire-format guide appended to the LLM’s system_instruction so it can parse the <ui_state> / <ui_event> messages. Defaults to UI_STATE_PROMPT_GUIDE; pass a string to override or None to disable. Living in system_instruction, it survives context resets.

property classifier: BaseClassifier

The classifier this worker asks the small questions about the screen.

property snapshot: dict[str, Any] | None

The latest accessibility snapshot of the page, or None before the first.

property selection: UISelection | None

The text the user has selected on the page, or None when nothing is.

async on_activated(args: dict | None) → None[source]

Set the classifier up with this worker’s task manager, then activate as usual.

Parameters:

args – Optional activation arguments.

async cleanup() → None[source]

Clean up the classifier along with the worker.

async send_command(name: str, payload: Any = None) → None[source]

Send a named UI command to the client.

Publishes a BusUICommandMessage; when RTVI is enabled, PipelineWorker translates it into an RTVIUICommandFrame on the pipeline. Client-side handlers subscribed to RTVIEvent.UICommand (or React’s useUICommandHandler) dispatch on the command name.

Parameters:
  • name – App-defined command name (e.g. "toast", "navigate", or any app-specific name).

  • payload –

    One of:

    • A pydantic BaseModel instance (including the built-in command models in pipecat.processors.frameworks.rtvi.models). Converted to a plain dict with model_dump().

    • A dataclass instance. Converted to a plain dict with dataclasses.asdict.

    • A dict forwarded as-is.

    • None, forwarded as an empty dict.

async scroll_to(ref: str) → None[source]

Send a scroll_to UI command to bring an element into view.

Convenience wrapper around send_command("scroll_to", ScrollTo(ref=ref)). These scroll_to / highlight / select_text / click / set_input_value helpers are plain methods, not LLM tools: compose them inside a @job handler or a custom @tool body.

Parameters:

ref – Snapshot ref (e.g. "e42") from the latest <ui_state>.

async highlight(ref: str) → None[source]

Send a highlight UI command to briefly flash an element.

Parameters:

ref – Snapshot ref (e.g. "e42") from the latest <ui_state>.

async select_text(ref: str, *, start_offset: int | None = None, end_offset: int | None = None) → None[source]

Send a select_text UI command to select an element’s text.

Selects the whole element by default, or the start_offset.. end_offset character sub-range (over the element’s concatenated textContent) when both are given. Used for deixis – pointing at content via the page’s text selection.

Parameters:
  • ref – Snapshot ref (e.g. "e42") from the latest <ui_state>.

  • start_offset – Optional start character offset of the selection.

  • end_offset – Optional end character offset (exclusive).

async click(ref: str) → None[source]

Send a click UI command (checkboxes, radios, submit buttons).

The standard client handler no-ops on disabled targets, so the worker can’t bypass affordances meant to be user-controlled.

Parameters:

ref – Snapshot ref (e.g. "e42") from the latest <ui_state>.

async set_input_value(ref: str, value: str, *, replace: bool = True) → None[source]

Send a set_input_value UI command to fill a text input/textarea.

Parameters:
  • ref – Snapshot ref (e.g. "e42") of the input or textarea.

  • value – Text to write into the field.

  • replace – When True (the default), overwrite the field; when False, append (e.g. to continue a long answer in a textarea).

async should_respond(message: BusUIEventMessage, criteria: str = 'the assistant should say something about what the user just did') → bool[source]

Ask the classifier whether a UI event calls for the assistant to speak.

Most clicks and edits need no comment, and a handler that reacts to events has to tell the few that do apart without an LLM turn. The question carries the event and the latest <ui_state> snapshot.

Parameters:
  • message – The UI event.

  • criteria – What is being checked for, as a yes or no question.

Returns:

Whether the assistant should respond to the event.

Raises:

ClassifierError – If the classifier could not answer.

async which_element(description: str, threshold: float = 0.5) → str | None[source]

Ask the classifier which element on screen the user means.

The candidates are the snapshot’s named elements, described by their role and name. The user’s words and the screen are the state.

Parameters:
  • description – What the user said, such as “the blue button”.

  • threshold – The probability below which no element is returned.

Returns:

The element’s snapshot ref, or None when there is no snapshot, no named element, or no confident answer.

Raises:

ClassifierError – If the classifier could not answer.

async check_screen(criteria: str) → YesNoResult[source]

Ask the classifier whether something is true of the screen.

Parameters:

criteria – What is being checked for, as a yes or no question, such as “is anything on the list still unchecked?”.

Returns:

How likely the answer is yes.

Raises:

ClassifierError – If the classifier could not answer.

async select_elements(criteria: str) → list[dict[str, Any]][source]

Ask the classifier which named elements on screen match a description.

One yes or no question per element, all in one call.

Parameters:

criteria – What the elements should be, such as “dairy products”.

Returns:

The matching elements, each as ref, label and probability, most likely first. Empty when nothing on screen matches or there is no snapshot.

Raises:

ClassifierError – If the classifier could not answer.

async act(action: str, description: str, *, value: str | None = None) → str | None[source]

Find the element the description means and act on it.

Parameters:
  • action – One of click, scroll_to, highlight, select_text or set_input_value.

  • description – The element in words, such as “the checkout button”.

  • value – The text to write, for set_input_value.

Returns:

The ref of the element acted on, or None when no element matched with enough confidence.

Raises:
list_elements(role: str | None = None) → list[dict[str, Any]][source]

The named elements on screen, without their refs.

Parameters:

role – Only elements of this role, such as checkbox; all when None.

Returns:

role, name, its state tags and, for an input, its value.

Return type:

One entry per element

async on_bus_message(message: BusMessage) → None[source]

Dispatch UI events alongside base lifecycle handling.

property current_job: BusJobRequestMessage | None

The job this worker is currently processing, or None when idle.

Set when a respond turn starts and cleared when the job completes. Lets @tool methods inspect the in-flight job without threading the message through every call.

Returns:

The in-flight BusJobRequestMessage, or None when idle.

render_query(message: BusJobRequestMessage) → str[source]

Extract the user’s query text from a job request.

Override to read a different payload shape. The returned string is appended to the LLM context as a user message before the LLM runs. The default reads payload["query"].

Parameters:

message – The inbound job request.

Returns:

The query text to feed into the LLM.

async respond_to_job(answer: str | None = None, *, tts_speak: bool = False, status: JobStatus = JobStatus.COMPLETED) → None[source]

Complete the in-flight job with the worker’s answer.

The reply the LLM writes completes the job on its own; call this from a @tool to answer with something else. The job responds with {"answer": answer} for the requester’s voice LLM to phrase, and a falsy answer completes the turn silently. No-op when no job is in flight or it was already answered.

Parameters:
  • answer – The worker’s answer, handed to the requester’s voice LLM to phrase.

  • tts_speak –

    Speak answer verbatim through the requester’s TTS and respond None.

    Deprecated since version 1.12.0: No replacement: respond with answer and let the requester’s voice LLM say it. Will be removed in 2.0.0.

  • status – Completion status. Defaults to JobStatus.COMPLETED.

async create_job_group_and_request_job(worker_names: list[str], **kwargs) → JobGroup[source]

Dispatch a job group and announce it to the client.

Parameters:
Returns:

The created JobGroup.

async cancel_job_group(job_id: str, *, reason: str | None = None) → None[source]

Cancel a running job group and complete its client card.

Parameters:
  • job_id – The job identifier to cancel.

  • reason – Optional human-readable reason for cancellation.

async on_job_update(message: BusJobUpdateMessage | BusJobUpdateUrgentMessage) → None[source]

Forward a worker’s progress update to the client.

async on_job_response(message: BusJobResponseMessage | BusJobResponseUrgentMessage) → None[source]

Forward a worker’s response to the client as its terminal envelope.

Runs before the group is torn down, so on an error status (with cancel_on_error) the client learns which worker failed before the card closes.

async on_job_stream_end(message: BusJobStreamEndMessage) → None[source]

Forward a worker’s stream end as its terminal envelope.

A worker may finish by ending its stream instead of responding; the client is told it completed, with the final stream data as the response payload.

async on_job_completed(result: JobGroupResponse) → None[source]

Complete the client’s card for a group whose workers all finished.

ui_job_group(*worker_names: str, name: str | None = None, payload: dict | None = None, timeout: float | None = None, cancel_on_error: bool = True, label: str | None = None, cancellable: bool = True) → JobGroupContext[source]

Deprecated wrapper for client-visible job groups.

Deprecated since version 1.8.0: Use job_group() instead, since every group a UIWorker dispatches is client-visible. Will be removed in 2.0.0.

async start_ui_job_group(*worker_names: str, name: str | None = None, payload: dict | None = None, timeout: float | None = None, cancel_on_error: bool = True, label: str | None = None, cancellable: bool = True) → str[source]

Deprecated wrapper for fire-and-forget client-visible job groups.

Deprecated since version 1.8.0: Use request_job_group() instead, since every group a UIWorker dispatches is client-visible. Will be removed in 2.0.0.

render_ui_state() → str[source]

Render the latest accessibility snapshot as a <ui_state> block.

Produces Playwright-MCP-style indented text with stable element refs. Apps inject the output via inject_ui_state() when they want the LLM to see what’s on screen.

When the snapshot carries a current text selection, a nested <selection ref="...">...</selection> block is appended inside <ui_state> so the LLM can resolve deictic references (“this paragraph”, “what I selected”) against on-page content.

Override to customize the rendered form.

Returns:

The <ui_state> block, or an empty string if no snapshot has been received yet.

async inject_ui_state() → None[source]

Append the latest <ui_state> block to the LLM context.

No-op when no snapshot has been received. Frame has run_llm=False — the snapshot is context, not a user turn.

render_ui_event(message: BusUIEventMessage) → str[source]

Render a UI event as a string for LLM context injection.

Override to customize the injected content. The default wraps the event in a single <ui_event> XML tag with a name attribute and a JSON-encoded payload as inner text.

Parameters:

message – The UI event to render.

Returns:

A string to append to the LLM context as a developer message.

pipecat.workers.ui.screen_tools(worker: str, *, timeout: float = 30.0) → list[source]

The tools that let a voice LLM ask a UIWorker about the screen, or act on it.

One tool, screen(action, target, value), sends the worker’s screen job and returns its answer as data, so the voice LLM learns what it asked and nothing more of the page. Hand them to the voice LLM’s context:

context = LLMContext(tools=screen_tools("ui"))
Parameters:
  • worker – The name of the UIWorker to ask.

  • timeout – Seconds to wait for an answer.

Returns:

The tools, as functions for an LLMContext.

pipecat.workers.ui.ui_event(name: str)[source]

Mark a worker method as a handler for a named UI event.

On UIWorker subclasses, decorated methods are automatically dispatched when a BusUIEventMessage with a matching name arrives.

Example:

class MyUIWorker(UIWorker):
    @ui_event("nav_click")
    async def on_nav(self, message):
        view = message.payload.get("view")
        ...
Parameters:

name – The UI event name to match.

Submodules