ui
UI worker: an LLM worker that observes and drives a client GUI over RTVI.
Composes the RTVI UI wire protocol (client events, accessibility snapshots,
server UI commands) with screen_tools, the tool a voice LLM uses to ask a
UIWorker about the screen. PipelineWorker connects a UIWorker to the
client automatically whenever RTVI is enabled — no decorator or separate
component to wire up.
- class pipecat.workers.ui.BaseUIWorker(**kwargs)[source]
Bases:
BaseWorkerWorker that surfaces its jobs and job groups on the client UI.
Deprecated since version 1.12.0: Use
UIWorkerinstead, which reports its job groups to the client itself. Will be removed in 2.0.0.Every group this worker dispatches is registered for lifecycle forwarding: a
group_startedenvelope is published at dispatch, worker updates and responses are forwarded asjob_update/job_completedenvelopes,group_completedis published at group teardown (normal completion, cancellation, or timeout), and the client’s reserved__cancel_job_groupevent is translated intocancel_job_groupfor groups dispatched as cancellable.Instantiable directly (no LLM): register one on the runner as a dispatcher when a pipeline app wants client-visible background work:
ui_jobs = BaseUIWorker("ui-jobs") job_id = await ui_jobs.request_job_group( "wikipedia", "news", params=JobGroupParams( payload={"query": query}, label=f"Research: {query}", ), )
- async create_job_group_and_request_job(worker_names: list[str], **kwargs) JobGroup[source]
Dispatch a job group and announce it to the client.
- Parameters:
worker_names – Names of the workers to send the job to.
**kwargs – Everything
create_job_group_and_request_job()takes, forwarded unchanged.
- Returns:
The created
JobGroup.
- async cancel_job_group(job_id: str, *, reason: str | None = None) None[source]
Cancel a running job group and complete its client card.
- Parameters:
job_id – The job identifier to cancel.
reason – Optional human-readable reason for cancellation.
- async on_bus_message(message: BusMessage) None[source]
Handle the client’s reserved
__cancel_job_groupevent.Everything else this worker forwards to the client hangs off the job hooks (
on_job_update(),on_job_response(),on_job_stream_end(),on_job_completed()), which the base class calls at the right point in a group’s lifecycle.- Parameters:
message – The
BusMessageto process.
- async on_job_update(message: BusJobUpdateMessage | BusJobUpdateUrgentMessage) None[source]
Forward a worker’s progress update to the client.
- async on_job_response(message: BusJobResponseMessage | BusJobResponseUrgentMessage) None[source]
Forward a worker’s response to the client as its terminal envelope.
Runs before the group is torn down, so on an error status (with
cancel_on_error) the client learns which worker failed before the card closes.
- async on_job_stream_end(message: BusJobStreamEndMessage) None[source]
Forward a worker’s stream end as its terminal envelope.
A worker may finish by ending its stream instead of responding; the client is told it completed, with the final stream data as the response payload.
- async on_job_completed(result: JobGroupResponse) None[source]
Complete the client’s card for a group whose workers all finished.
- class pipecat.workers.ui.BusUICommandMessage(command_name: str = '', payload: Any = None, *, source: str, target: str | None = None)[source]
Bases:
BusUIDataMessageA UI command sent from a server-side worker to the client.
Published by
UIWorker.send_command(name, payload). The client-facing worker translates it into a command on its client’s wire format; over RTVI that is anRTVIUICommandFrame(command=command_name, payload=payload)pushed through the pipeline.- Parameters:
command_name – App-defined command name.
payload – App-defined payload (already a plain dict by the time it lands on the bus).
- class pipecat.workers.ui.BusUIEventMessage(event_name: str = '', payload: Any = None, *, source: str, target: str | None = None)[source]
Bases:
BusUIDataMessageA UI event sent from the client to a server-side worker.
Emitted by the client-facing worker when the client dispatches an event via
PipecatClient.sendUIEvent(event, payload).UIWorkersubclasses dispatch these to@ui_event(name)handlers.- Parameters:
event_name – App-defined event name.
payload – App-defined payload. Schemaless by design.
- class pipecat.workers.ui.BusUIJobCompletedMessage(job_id: str = '', worker_name: str = '', status: str = '', response: Any = None, at: int = 0, *, source: str, target: str | None = None)[source]
Bases:
BusUIDataMessageA worker in a user-facing job group has completed.
Forwarded by a
UIWorkerwhen a worker of one of its job groups reaches a terminal state. The client-facing worker forwards it as aui-job-groupenvelope withkind = "job_completed".- Parameters:
job_id – The shared job-group identifier.
worker_name – The worker that produced the response.
status – Completion status as a string (
JobStatusvalue).response – The worker’s response payload.
at – Epoch milliseconds when the response was received.
- class pipecat.workers.ui.BusUIJobGroupCompletedMessage(job_id: str = '', at: int = 0, *, source: str, target: str | None = None)[source]
Bases:
BusUIDataMessageA user-facing job group has completed.
Published by a
UIWorkeronce every worker in the group has finished, or the group was cancelled. The client-facing worker forwards it as aui-job-groupenvelope withkind = "group_completed".- Parameters:
job_id – The shared job-group identifier.
at – Epoch milliseconds when the group completed.
- class pipecat.workers.ui.BusUIJobGroupStartedMessage(job_id: str = '', workers: list[str] | None = None, label: str | None = None, cancellable: bool = True, at: int = 0, *, source: str, target: str | None = None)[source]
Bases:
BusUIDataMessageA user-facing job group has been dispatched.
Published by a
UIWorkeras it dispatches the group. The client-facing worker forwards it as aui-job-groupenvelope withkind = "group_started".- Parameters:
job_id – Shared job-group identifier for the group.
workers – Names of the workers the work was dispatched to.
label – Optional human-readable label for the group.
cancellable – Whether the client may request cancellation.
at – Epoch milliseconds when the group started.
- class pipecat.workers.ui.BusUIJobUpdateMessage(job_id: str = '', worker_name: str = '', data: Any = None, at: int = 0, *, source: str, target: str | None = None)[source]
Bases:
BusUIDataMessagePer-worker progress for a user-facing job group.
Forwarded by a
UIWorkerwhenever a worker of one of its job groups emits aBusJobUpdateMessage. The client-facing worker forwards it as aui-job-groupenvelope withkind = "job_update".- Parameters:
job_id – The shared job-group identifier.
worker_name – The worker that produced the update.
data – The worker’s update payload, forwarded verbatim.
at – Epoch milliseconds when the update was emitted on the bus.
- class pipecat.workers.ui.ReplyToolMixin(**kwargs)[source]
Bases:
objectExpose a
replytool covering the full standard action set.Deprecated since version 1.12.0: Use
screen_tools()instead: the voice LLM asks the worker through thescreenjob and says the answer itself. Will be removed in 2.0.0.Single bundled LLM tool with a required spoken
answerplus optional visual and state-changing actions. One tool call per turn, no chaining; the requiredanswerargument is enforced by the API schema so the model cannot omit the terminator.Compose alongside
UIWorker:class MyUIWorker(ReplyToolMixin, UIWorker): ...
Covers pointing apps (
scroll_to+highlight), reading apps (scroll_to+select_text), form apps (fills+click), and any blend (e.g. a document review with selection-based deixis AND voice-driven note-taking). The LLM uses whichever fields fit the user’s request per turn; unused fields staynulland don’t affect behavior.Delivers
answeras verbatim TTS (respond_to_job(answer, tts_speak=True)) – the worker speaks the exact phrase. Apps that want a minimal schema (only the fields actually used, or app-specific commands), or that want the requester’s voice LLM to phrase the reply instead, write their own@tool replyon theUIWorkersubclass directly. Use the helper methods onUIWorkerplussend_commandto dispatch the underlying UI commands.The host class must provide
scroll_to,highlight,select_text,click,set_input_value, andrespond_to_job(UIWorkerdoes) and must be the target of@tooldiscovery on the LLM pipeline.- async reply(params: FunctionCallParams, answer: str, scroll_to: str | None = None, highlight: list[str] | None = None, select_text: str | None = None, fills: list[dict] | None = None, click: list[str] | None = None)[source]
Reply to the user. Optionally point at content and act on inputs.
Always called exactly once per turn.
answeris required; the action fields are optional and may be combined.Visual / pointing actions (draw the user’s attention):
scroll_tobrings an element into view (single ref).highlightflashes elements briefly (list of refs). Best for short emphasis like a button or a fact.select_textputs the page’s text selection on an element (single ref). Best for “this paragraph” / “the section about X” so the user sees exactly what was meant. Persists until the user clicks elsewhere.
State-changing actions (modify form / app state):
fillswrites values into inputs (list of{"ref", "value"}objects, multi-fill in one turn).clickclicks elements (list of refs in order). Use for checkboxes, radios, submit buttons.
Order of dispatch within a turn:
scroll_to, thenhighlight, thenselect_text, thenfills, thenclick, then speak the answer.- Parameters:
params – Framework-provided tool invocation context.
answer – The spoken reply in plain language. One short sentence. No markdown, no symbols.
scroll_to – Optional snapshot ref. Scrolls the element into view before speaking.
highlight – Optional list of snapshot refs. Visually pulses each element.
select_text – Optional snapshot ref. Places the page’s text selection on that element.
fills – Optional list of
{"ref": "eN", "value": "..."}objects. Writes each value into the input atref.click – Optional list of snapshot refs to click in order.
- class pipecat.workers.ui.UISelection(ref: str, text: str)[source]
Bases:
NamedTupleThe text the user has selected on the page, and the element it is in.
- class pipecat.workers.ui.UIWorker(name: str, *, llm: LLMService[Any], context: LLMContext | None = None, classifier: BaseClassifier | None = None, assistant_params: LLMAssistantAggregatorParams | None = None, inject_events: bool = True, auto_inject_ui_state: bool = True, keep_history: bool = False, prompt_guide: str | None = '## UI context\n\nYour developer context includes two kinds of SDK-managed messages:\n\n- ``<ui_event name="..." >payload</ui_event>``: an event the user just triggered on the client (click, tab switch, navigation, etc.). The payload is JSON for that event.\n- ``<ui_state>...</ui_state>``: an accessibility snapshot of the current screen, injected at the start of every turn. Indented tree in Playwright-MCP style. Each line is ``- role "name" [state] [ref=eN]`` with children nested one level deeper. A line can also carry ``= "value"`` (an element\'s current value, e.g. text already typed into an input) and ``[level=N]`` (heading depth).\n\nState tags include ``[focused]``, ``[selected]``, ``[disabled]``, and ``[offscreen]``. A node tagged ``[offscreen]`` exists on the page but is not currently in the user\'s viewport; only visible (non-offscreen) nodes count for position-based references.\n\nGrids carry a ``[cols=N]`` tag. Their cells are listed in reading order (left-to-right, top-to-bottom); with N columns, cell K sits at row ``ceil(K/N)``, column ``((K-1) mod N) + 1``. Example with ``[cols=8]`` and 16 children: "top right" is cell 8, "bottom left" is cell 9.\n\nResolve position references ("top right", "the first one", "the third new release") against the most recent ``<ui_state>`` tree. Sibling order matches reading order on screen (top-to-bottom, left-to-right within each region).\n\nWhen the user has text selected on the page, the snapshot ends with a ``<selection ref="eN">selected text</selection>`` block inside ``<ui_state>``. Treat the selection as the deictic referent for "this", "that", "what I selected", and similar phrases. The ``ref`` identifies the closest enclosing element that has a ref in the tree; the inner text is the actual selected content (truncated if very long). Text inside ``<input>`` or ``<textarea>`` selections is faithful to ``selectionStart``/``selectionEnd`` on the element.\n\nRefs (``e42``) are stable handles for acting on elements: pass the ``ref`` from the most recent ``<ui_state>`` to any tool that operates on a node. The same element keeps its ref across snapshots while it stays on the page, so you can refer back to it across turns. Always resolve refs against the latest snapshot, and bring an ``[offscreen]`` element into view before acting on it.')[source]
Bases:
LLMContextWorkerLLM worker that reads and drives a client GUI over the RTVI UI channel.
A
UIWorkerconnects an LLM to whatever the user is looking at: it sees the screen as accessibility snapshots, reacts to the user’s UI events, and acts on the page by sending commands to the client. It is the delegate side of a voice/UI split – a voice layer (the main pipeline’s LLM, or a separateLLMWorker) handles speech and hands screen-relevant work to this worker.Capabilities:
See the screen. The latest accessibility snapshot is rendered as
<ui_state>and auto-injected into the LLM context before each inference. Code reads it throughsnapshotand the user’sselection.React to UI events, dispatched to
@ui_event(name)handlers.Drive the UI with
send_commandand thescroll_to/highlight/select_text/click/set_input_valuehelpers.Decide small things with a classifier, not an LLM turn: whether a UI event deserves a comment (
should_respond), which element on screen the user means (which_element), whether something is true of the screen (check_screen), which elements match a description (select_elements), andacton an element named in words.Answer a voice LLM’s questions about the screen through the
screenjob: find an element, check whether something is true, select the elements matching a description, list what is on screen, read the user’s selection, or click, scroll to, highlight, select or fill an element. Every answer is short data and never the page;screen_tools()gives the voice LLM the tool that sends it.Answer as a delegate. The built-in single-flight
respondjob runs one screen-grounded LLM turn and answers with the reply the LLM writes. A@toolthat callsrespond_to_jobanswers instead when it needs to decide how the answer reaches the user.Surface long work. Every job group this worker dispatches is reported to the client as it goes: a card when the group starts, a line per worker’s progress and completion, and the close when the group completes, whether normally, by cancellation or by timeout. The client can cancel a group dispatched as cancellable.
PipelineWorkerconnects a UIWorker to the client automatically when RTVI is enabled – no extra wiring. A working worker needs only an LLM; subclass it to react to UI events or add jobs and tools, and overriderender_queryto read a non-default job payload.Example:
class MyUIWorker(UIWorker): @ui_event("nav_click") async def on_nav(self, message): view = message.payload.get("view") ... worker = MyUIWorker("ui", llm=OpenAILLMService(api_key="..."))
Note
With client
trackViewporton (the default), off-screen nodes carry[offscreen]in<ui_state>;scroll_tobefore acting on them.- __init__(name: str, *, llm: LLMService[Any], context: LLMContext | None = None, classifier: BaseClassifier | None = None, assistant_params: LLMAssistantAggregatorParams | None = None, inject_events: bool = True, auto_inject_ui_state: bool = True, keep_history: bool = False, prompt_guide: str | None = '## UI context\n\nYour developer context includes two kinds of SDK-managed messages:\n\n- ``<ui_event name="..." >payload</ui_event>``: an event the user just triggered on the client (click, tab switch, navigation, etc.). The payload is JSON for that event.\n- ``<ui_state>...</ui_state>``: an accessibility snapshot of the current screen, injected at the start of every turn. Indented tree in Playwright-MCP style. Each line is ``- role "name" [state] [ref=eN]`` with children nested one level deeper. A line can also carry ``= "value"`` (an element\'s current value, e.g. text already typed into an input) and ``[level=N]`` (heading depth).\n\nState tags include ``[focused]``, ``[selected]``, ``[disabled]``, and ``[offscreen]``. A node tagged ``[offscreen]`` exists on the page but is not currently in the user\'s viewport; only visible (non-offscreen) nodes count for position-based references.\n\nGrids carry a ``[cols=N]`` tag. Their cells are listed in reading order (left-to-right, top-to-bottom); with N columns, cell K sits at row ``ceil(K/N)``, column ``((K-1) mod N) + 1``. Example with ``[cols=8]`` and 16 children: "top right" is cell 8, "bottom left" is cell 9.\n\nResolve position references ("top right", "the first one", "the third new release") against the most recent ``<ui_state>`` tree. Sibling order matches reading order on screen (top-to-bottom, left-to-right within each region).\n\nWhen the user has text selected on the page, the snapshot ends with a ``<selection ref="eN">selected text</selection>`` block inside ``<ui_state>``. Treat the selection as the deictic referent for "this", "that", "what I selected", and similar phrases. The ``ref`` identifies the closest enclosing element that has a ref in the tree; the inner text is the actual selected content (truncated if very long). Text inside ``<input>`` or ``<textarea>`` selections is faithful to ``selectionStart``/``selectionEnd`` on the element.\n\nRefs (``e42``) are stable handles for acting on elements: pass the ``ref`` from the most recent ``<ui_state>`` to any tool that operates on a node. The same element keeps its ref across snapshots while it stays on the page, so you can refer back to it across turns. Always resolve refs against the latest snapshot, and bring an ``[offscreen]`` element into view before acting on it.')[source]
Initialize the UIWorker.
- Parameters:
name – Unique name for this worker.
llm – The LLM service.
context – Optional pre-built
LLMContext. Seeded messages are part of the mutable history and are cleared on eachkeep_history=Falsereset; put durable instructions in the LLM’ssystem_instructioninstead.classifier – Answers the small questions about the screen (
should_respond,which_element). Without one, thellmanswers them through anLLMClassifier, which costs an LLM call per question; aJevClassifieranswers in about a tenth of a second with a calibrated probability.assistant_params – Optional assistant-aggregator parameters, e.g. to enable context summarization for
keep_history=Trueworkers.inject_events – When True (the default), append each UI event to the context as a
<ui_event>developer message. Overriderender_ui_eventto change the content, or set False to disable.auto_inject_ui_state – When True (the default), append the latest
<ui_state>snapshot to the context before every inference (via the LLM’son_before_process_framehook). Set False to inject manually withinject_ui_state().keep_history – When False (the default), the context is cleared at the start of every job, so each turn sees only the current
<ui_state>and query – best for the stateless-delegate role. When True, history accumulates across jobs so the LLM can resolve multi-turn references (“the next one”, “the Pro version”), at the cost of more tokens and possible confusion from stale<ui_state>blocks. Use context summarization to prune the history when it gets too large.prompt_guide – Wire-format guide appended to the LLM’s
system_instructionso it can parse the<ui_state>/<ui_event>messages. Defaults toUI_STATE_PROMPT_GUIDE; pass a string to override orNoneto disable. Living insystem_instruction, it survives context resets.
- property classifier: BaseClassifier
The classifier this worker asks the small questions about the screen.
- property snapshot: dict[str, Any] | None
The latest accessibility snapshot of the page, or
Nonebefore the first.
- property selection: UISelection | None
The text the user has selected on the page, or
Nonewhen nothing is.
- async on_activated(args: dict | None) None[source]
Set the classifier up with this worker’s task manager, then activate as usual.
- Parameters:
args – Optional activation arguments.
- async send_command(name: str, payload: Any = None) None[source]
Send a named UI command to the client.
Publishes a
BusUICommandMessage; when RTVI is enabled,PipelineWorkertranslates it into anRTVIUICommandFrameon the pipeline. Client-side handlers subscribed toRTVIEvent.UICommand(or React’suseUICommandHandler) dispatch on the command name.- Parameters:
name – App-defined command name (e.g.
"toast","navigate", or any app-specific name).payload –
One of:
A pydantic
BaseModelinstance (including the built-in command models inpipecat.processors.frameworks.rtvi.models). Converted to a plain dict withmodel_dump().A dataclass instance. Converted to a plain dict with
dataclasses.asdict.A
dictforwarded as-is.None, forwarded as an empty dict.
- async scroll_to(ref: str) None[source]
Send a
scroll_toUI command to bring an element into view.Convenience wrapper around
send_command("scroll_to", ScrollTo(ref=ref)). Thesescroll_to/highlight/select_text/click/set_input_valuehelpers are plain methods, not LLM tools: compose them inside a@jobhandler or a custom@toolbody.- Parameters:
ref – Snapshot ref (e.g.
"e42") from the latest<ui_state>.
- async highlight(ref: str) None[source]
Send a
highlightUI command to briefly flash an element.- Parameters:
ref – Snapshot ref (e.g.
"e42") from the latest<ui_state>.
- async select_text(ref: str, *, start_offset: int | None = None, end_offset: int | None = None) None[source]
Send a
select_textUI command to select an element’s text.Selects the whole element by default, or the
start_offset..end_offsetcharacter sub-range (over the element’s concatenatedtextContent) when both are given. Used for deixis – pointing at content via the page’s text selection.- Parameters:
ref – Snapshot ref (e.g.
"e42") from the latest<ui_state>.start_offset – Optional start character offset of the selection.
end_offset – Optional end character offset (exclusive).
- async click(ref: str) None[source]
Send a
clickUI command (checkboxes, radios, submit buttons).The standard client handler no-ops on
disabledtargets, so the worker can’t bypass affordances meant to be user-controlled.- Parameters:
ref – Snapshot ref (e.g.
"e42") from the latest<ui_state>.
- async set_input_value(ref: str, value: str, *, replace: bool = True) None[source]
Send a
set_input_valueUI command to fill a text input/textarea.- Parameters:
ref – Snapshot ref (e.g.
"e42") of the input or textarea.value – Text to write into the field.
replace – When True (the default), overwrite the field; when False, append (e.g. to continue a long answer in a textarea).
- async should_respond(message: BusUIEventMessage, criteria: str = 'the assistant should say something about what the user just did') bool[source]
Ask the classifier whether a UI event calls for the assistant to speak.
Most clicks and edits need no comment, and a handler that reacts to events has to tell the few that do apart without an LLM turn. The question carries the event and the latest
<ui_state>snapshot.- Parameters:
message – The UI event.
criteria – What is being checked for, as a yes or no question.
- Returns:
Whether the assistant should respond to the event.
- Raises:
ClassifierError – If the classifier could not answer.
- async which_element(description: str, threshold: float = 0.5) str | None[source]
Ask the classifier which element on screen the user means.
The candidates are the snapshot’s named elements, described by their role and name. The user’s words and the screen are the state.
- Parameters:
description – What the user said, such as “the blue button”.
threshold – The probability below which no element is returned.
- Returns:
The element’s snapshot ref, or
Nonewhen there is no snapshot, no named element, or no confident answer.- Raises:
ClassifierError – If the classifier could not answer.
- async check_screen(criteria: str) YesNoResult[source]
Ask the classifier whether something is true of the screen.
- Parameters:
criteria – What is being checked for, as a yes or no question, such as “is anything on the list still unchecked?”.
- Returns:
How likely the answer is yes.
- Raises:
ClassifierError – If the classifier could not answer.
- async select_elements(criteria: str) list[dict[str, Any]][source]
Ask the classifier which named elements on screen match a description.
One yes or no question per element, all in one call.
- Parameters:
criteria – What the elements should be, such as “dairy products”.
- Returns:
The matching elements, each as
ref,labelandprobability, most likely first. Empty when nothing on screen matches or there is no snapshot.- Raises:
ClassifierError – If the classifier could not answer.
- async act(action: str, description: str, *, value: str | None = None) str | None[source]
Find the element the description means and act on it.
- Parameters:
action – One of
click,scroll_to,highlight,select_textorset_input_value.description – The element in words, such as “the checkout button”.
value – The text to write, for
set_input_value.
- Returns:
The ref of the element acted on, or
Nonewhen no element matched with enough confidence.- Raises:
ValueError – If
actionis not one of the five.ClassifierError – If the classifier could not answer.
- list_elements(role: str | None = None) list[dict[str, Any]][source]
The named elements on screen, without their refs.
- Parameters:
role – Only elements of this role, such as
checkbox; all whenNone.- Returns:
role,name, itsstatetags and, for an input, itsvalue.- Return type:
One entry per element
- async on_bus_message(message: BusMessage) None[source]
Dispatch UI events alongside base lifecycle handling.
- property current_job: BusJobRequestMessage | None
The job this worker is currently processing, or
Nonewhen idle.Set when a respond turn starts and cleared when the job completes. Lets
@toolmethods inspect the in-flight job without threading the message through every call.- Returns:
The in-flight
BusJobRequestMessage, orNonewhen idle.
- render_query(message: BusJobRequestMessage) str[source]
Extract the user’s query text from a job request.
Override to read a different payload shape. The returned string is appended to the LLM context as a user message before the LLM runs. The default reads
payload["query"].- Parameters:
message – The inbound job request.
- Returns:
The query text to feed into the LLM.
- async respond_to_job(answer: str | None = None, *, tts_speak: bool = False, status: JobStatus = JobStatus.COMPLETED) None[source]
Complete the in-flight job with the worker’s answer.
The reply the LLM writes completes the job on its own; call this from a
@toolto answer with something else. The job responds with{"answer": answer}for the requester’s voice LLM to phrase, and a falsyanswercompletes the turn silently. No-op when no job is in flight or it was already answered.- Parameters:
answer – The worker’s answer, handed to the requester’s voice LLM to phrase.
tts_speak –
Speak
answerverbatim through the requester’s TTS and respondNone.Deprecated since version 1.12.0: No replacement: respond with
answerand let the requester’s voice LLM say it. Will be removed in 2.0.0.status – Completion status. Defaults to
JobStatus.COMPLETED.
- async create_job_group_and_request_job(worker_names: list[str], **kwargs) JobGroup[source]
Dispatch a job group and announce it to the client.
- Parameters:
worker_names – Names of the workers to send the job to.
**kwargs – Everything
create_job_group_and_request_job()takes, forwarded unchanged.
- Returns:
The created
JobGroup.
- async cancel_job_group(job_id: str, *, reason: str | None = None) None[source]
Cancel a running job group and complete its client card.
- Parameters:
job_id – The job identifier to cancel.
reason – Optional human-readable reason for cancellation.
- async on_job_update(message: BusJobUpdateMessage | BusJobUpdateUrgentMessage) None[source]
Forward a worker’s progress update to the client.
- async on_job_response(message: BusJobResponseMessage | BusJobResponseUrgentMessage) None[source]
Forward a worker’s response to the client as its terminal envelope.
Runs before the group is torn down, so on an error status (with
cancel_on_error) the client learns which worker failed before the card closes.
- async on_job_stream_end(message: BusJobStreamEndMessage) None[source]
Forward a worker’s stream end as its terminal envelope.
A worker may finish by ending its stream instead of responding; the client is told it completed, with the final stream data as the response payload.
- async on_job_completed(result: JobGroupResponse) None[source]
Complete the client’s card for a group whose workers all finished.
- ui_job_group(*worker_names: str, name: str | None = None, payload: dict | None = None, timeout: float | None = None, cancel_on_error: bool = True, label: str | None = None, cancellable: bool = True) JobGroupContext[source]
Deprecated wrapper for client-visible job groups.
Deprecated since version 1.8.0: Use
job_group()instead, since every group aUIWorkerdispatches is client-visible. Will be removed in 2.0.0.
- async start_ui_job_group(*worker_names: str, name: str | None = None, payload: dict | None = None, timeout: float | None = None, cancel_on_error: bool = True, label: str | None = None, cancellable: bool = True) str[source]
Deprecated wrapper for fire-and-forget client-visible job groups.
Deprecated since version 1.8.0: Use
request_job_group()instead, since every group aUIWorkerdispatches is client-visible. Will be removed in 2.0.0.
- render_ui_state() str[source]
Render the latest accessibility snapshot as a
<ui_state>block.Produces Playwright-MCP-style indented text with stable element refs. Apps inject the output via
inject_ui_state()when they want the LLM to see what’s on screen.When the snapshot carries a current text selection, a nested
<selection ref="...">...</selection>block is appended inside<ui_state>so the LLM can resolve deictic references (“this paragraph”, “what I selected”) against on-page content.Override to customize the rendered form.
- Returns:
The
<ui_state>block, or an empty string if no snapshot has been received yet.
- async inject_ui_state() None[source]
Append the latest
<ui_state>block to the LLM context.No-op when no snapshot has been received. Frame has
run_llm=False— the snapshot is context, not a user turn.
- render_ui_event(message: BusUIEventMessage) str[source]
Render a UI event as a string for LLM context injection.
Override to customize the injected content. The default wraps the event in a single
<ui_event>XML tag with anameattribute and a JSON-encoded payload as inner text.- Parameters:
message – The UI event to render.
- Returns:
A string to append to the LLM context as a developer message.
- pipecat.workers.ui.screen_tools(worker: str, *, timeout: float = 30.0) list[source]
The tools that let a voice LLM ask a
UIWorkerabout the screen, or act on it.One tool,
screen(action, target, value), sends the worker’sscreenjob and returns its answer as data, so the voice LLM learns what it asked and nothing more of the page. Hand them to the voice LLM’s context:context = LLMContext(tools=screen_tools("ui"))
- Parameters:
worker – The name of the
UIWorkerto ask.timeout – Seconds to wait for an answer.
- Returns:
The tools, as functions for an
LLMContext.
- pipecat.workers.ui.ui_event(name: str)[source]
Mark a worker method as a handler for a named UI event.
On
UIWorkersubclasses, decorated methods are automatically dispatched when aBusUIEventMessagewith a matchingnamearrives.Example:
class MyUIWorker(UIWorker): @ui_event("nav_click") async def on_nav(self, message): view = message.payload.get("view") ...
- Parameters:
name – The UI event name to match.
Submodules
- ui_event_decorator
- ui_job_context
- ui_prompts
- ui_tools
- ui_worker
UISelectionUIWorkerUIWorker.__init__()UIWorker.classifierUIWorker.snapshotUIWorker.selectionUIWorker.on_activated()UIWorker.cleanup()UIWorker.send_command()UIWorker.scroll_to()UIWorker.highlight()UIWorker.select_text()UIWorker.click()UIWorker.set_input_value()UIWorker.should_respond()UIWorker.which_element()UIWorker.check_screen()UIWorker.select_elements()UIWorker.act()UIWorker.list_elements()UIWorker.on_bus_message()UIWorker.current_jobUIWorker.render_query()UIWorker.respond_to_job()UIWorker.create_job_group_and_request_job()UIWorker.cancel_job_group()UIWorker.on_job_update()UIWorker.on_job_response()UIWorker.on_job_stream_end()UIWorker.on_job_completed()UIWorker.ui_job_group()UIWorker.start_ui_job_group()UIWorker.render_ui_state()UIWorker.inject_ui_state()UIWorker.render_ui_event()