Skip to content

YELP 2026-09-29 Tier 3

Read original ↗

Yelp — The making of the Yelp Assistant UI

Summary

Yelp Engineering post (2026-09-29) on the UI architecture behind Yelp Assistant — Yelp's first large-scale consumer-facing LLM product (shipped early 2024 on iOS), which connects users with local service pros ("I have a leaky faucet" → the assistant asks clarifying questions → submits a Request-a-Quote project to multiple businesses) and has since grown to answer business questions, make reservations, book appointments, and order food. The central design problem is decision reversibility under uncertainty: the team was learning to use LLMs as a product and the technology was moving fast, so they wanted an architecture where a UI decision could be made — and un-made — cheaply. Mobile's app-store constraint (any change takes days to reach users, and some users never update) pushed them toward server-driven UI (SDUI) via Yelp's CHAOS framework, where a backend change reaches all clients within minutes. But CHAOS (then in its infancy) lacked network-request affordances and text-input support, so they chose a hybrid: the top bar and input bar are native, the scrolling chat content is server-driven. The post details how the two halves are stitched together — view/data separation (fetch the CHAOS view config once at launch as a template, then hydrate it turn-by-turn with data), a turn-based reactive chat loop with optimistic updates (insert the user's message + a typing indicator locally the instant Send is tapped; the client owns inserting/removing the typing indicator), a small vocabulary of MessageEvent types (user_message, typing_indicator, bot_message, link_message, quick_replies), CHAOS actions for user-triggered behavior (quick_reply, open_url), and client actions for backend-triggered behavior (DisableChat, sent in the bot's final response to remove the input bar and end the session). Recoverable vs non-recoverable failures degrade gracefully. No operational numbers — this is an architecture walkthrough, not a retrospective.

Key takeaways

  1. Reversible decisions were the explicit design goal. Verbatim: "We needed to be able to make decisions and feel comfortable reversing them when they no longer made sense. If decisions are paid with time, how can we mitigate the cost of making and changing decisions?" The whole architecture is organized around making UI changes cheap to ship and cheap to undo — which is exactly what SDUI buys on mobile. (Source: sources/2026-09-29-yelp-the-making-of-the-yelp-assistant-ui)

  2. The native-vs-SDUI choice is a mobile-specific deployment-latency argument. "Unlike building for web platforms where users are one refresh away from using the latest client code, mobile applications have a unique constraint: every change requires a new build to be deployed to and reviewed by the platform's app store … any change, however small, can take several days to reach the users in the best of cases. In the worst cases, a user may never update the app, and they will remain in a time-locked experience." SDUI collapses that to "update the backend and all clients receive the latest configuration within minutes." The team "litigated, deliberated, decided, and backtracked multiple times" before committing. Canonical SDUI motivation.

  3. They went hybrid because SDUI didn't cover everything — ship-and-learn beat build-the-missing-primitives. CHAOS "shines at rendering dynamic content such as a conversation" but "doesn't provide affordances to send requests to the backend and support for text input fields did not exist at the time." Rather than invest in building those into CHAOS, they split the screen: top bar + input bar native, chat content server-driven. "Our goal was to build something, ship, and start learning." This is the load-bearing hybrid-native-sdui decision — SDUI where it's strong (dynamic list of heterogeneous messages), native where it's weak (text entry, network calls).

  4. View and data are separated: fetch the template once, hydrate per turn. "the chosen architecture separates the view and the data. The client fetches the view and uses it as a template, and as the conversation progresses, the data hydrates the view." At launch the app fetches the CHAOS view configuration (the blueprint + the actions that fire on certain events + operations needed for correct functioning), then sets up datasets (CHAOS's client-side state store); "each entry in the dataset represents a MessageEvent." CHAOS observes the datasets and keeps the UI current as messages arrive. This view-data-separation is what lets new content be appended incrementally without re-fetching the view each turn — separation of concerns applied to the client/server contract: "Whether the client renders a user message, a typing indicator or something completely new does not matter, as CHAOS abstracts it away."

  5. The chat loop is turn-based, reactive, and uses optimistic updates. Yelp Assistant "is reactive and turn-based. It only responds to a message and doesn't proactively send messages … the user must wait for the bot to reply or fail before sending a new message." On Send: (1) disable the input box, (2) clear quick replies, (3) send the message to the backend, (4) optimistically insert the user message + a typing indicator into the dataset. On reply: remove the typing indicator, insert the response, add quick replies if any, re-enable input. "We insert the user message to the dataset to show it instantly. We call this an optimistic update." "The client is fully responsible for inserting and removing the typing indicator at the right time."

  6. MessageEvent is the opaque unit of the conversation; a small typed vocabulary drives rendering. "Every message in the conversation is a MessageEvent and to the chat application, it is mostly opaque … CHAOS uses component templates to render each one." Types disclosed (with ephemeral/persistent + interactive flags):

Type Ephemeral/Persistent Interactive Notes
user_message Persistent No The user's text
typing_indicator Ephemeral No Shown from Send until response
bot_message Persistent No The bot's reply
link_message Persistent Yes Navigate to another screen; generally signals chat end
quick_replies Ephemeral Yes Tappable suggested replies

Only quick_replies and typing_indicator are removed from the dataset; quick_replies are backed by a different dataset but follow the same principle. The MessageEvent shape is minimal — data class MessageEvent(id, type, text?, url?).

  1. Two kinds of actions: CHAOS actions (user-triggered) vs client actions (backend-triggered). CHAOS actions are declared in event hooks (onView, onClick) and fire client code when the user interacts — quick_reply (a Yelp Assistant-specific action that routes a tapped bubble back through the message-sending logic) and open_url (built into CHAOS; navigates to the link). Client actions invert the direction: "a client action is executed by the client on the backend's instruction." The first client action, DisableChat, is sent in the bot's last response to "disable all forms of input and prevent further messages" — most commonly because the project was submitted, but also on rate limits or errors. "The mechanism is flexible to support new actions as we require them." This backend-driven-client-action seam is how the server ends a session or otherwise steers a native-owned surface it doesn't directly render.

  2. Failures degrade gracefully in two tiers. "We handle two failure modes categorized as recoverable or non-recoverable. In both cases, the client removes the typing indicator and inserts an error message." Recoverable: retry a bounded number of times, then ask the user to resend. Non-recoverable: ask the user to try again later and disable the chat, ending the session. Canonical graceful degradation applied to an LLM chat turn.

  3. The hybrid + view/data split paid a real reuse dividend. The architecture "makes it possible to add new functionality such as image recognition or support for logged-out users as each part of the entire system is independent of the others." And it seeded a lineage: "a team cloned and repurposed Android's Yelp Assistant to bootstrap the pilot for Biz Ask Anything … We didn't realize it at the time, but we laid the groundwork for what would become the next generation of Yelp Assistant." Costs are named too: onboarding time (people must learn Yelp Assistant + CHAOS before contributing) and cross-platform consistency bugs ("what works well in one platform is buggy in another").

Systems

  • Yelp Assistant — the product. This post finally documents the UI architecture of the original consumer/services-pro variant (the 2026-03 BAA post had only named it). Hybrid native-shell + SDUI-content chat app.
  • CHAOS — Yelp's SDUI framework; here used in hybrid mode (content only) with product-specific extensions: custom components, custom CHAOS actions (quick_reply), client actions (DisableChat), datasets as the client-side conversation store, and MessageEvent as the per-message data unit.

Concepts

  • server-driven UI (SDUI) — the backend tells the client what to show and what to do on interaction; a backend deploy reaches all clients in minutes, bypassing app-store latency. This post is a canonical motivation + a canonical hybrid instance.
  • optimistic update — insert the user message + typing indicator into local state immediately on Send, before the backend confirms; reconcile (remove typing indicator, insert reply) when the response arrives, or roll into an error message on failure.
  • separation of concerns — view (CHAOS template, fetched once) vs data (MessageEvents, streamed per turn); the renderer is agnostic to message type.
  • graceful degradation — recoverable (bounded retry → ask to resend) vs non-recoverable (ask-later + disable chat) failure tiers.
  • client-server model — the backend as an actor that can push behavior to the client (client actions), not just serve data on request.

Operational numbers

None disclosed. No latency, RPS, message-volume, or platform-distribution figures — the post is an architecture walkthrough. First version launched early 2024 on iOS (SwiftUI); later expanded to Android and web via the hybrid CHAOS approach.

Caveats

  • Architecture walkthrough, not a retrospective; no metrics.
  • The code snippets (MessageEvent, QuickReply, Response) are illustrative Kotlin-flavored pseudocode, not verbatim production types.
  • "The specifics of how CHAOS achieves" per-MessageEvent rendering "will be covered in a future blog post" — the CHAOS-internal template mechanism is out of scope here (see the 2025-07-08 CHAOS backend post for the general model).
  • The named event types are "a few event types to support the experience we were building" — an evolving, not exhaustive, vocabulary.

Source

Last updated · 766 distilled / 2,225 read