Skip to content

YELP 2026-08-20

Read original ↗

Building Menu Vision: Real-Time Dish Recognition

Summary

Yelp's Menu Vision is a mobile feature that turns a phone camera pointed at a restaurant menu into an augmented-reality dish-discovery surface: it runs on-device text recognition over live camera frames, matches the recognized menu text against Yelp's curated per-business dish data, and overlays each recognized dish with its Yelp photos, review counts, price, and a "Popular" badge. The post is a from-hackathon-to-production narrative, but it carries real system-design content: a precompute-to-Cassandra data pipeline that combines multiple menu sources (partner/owner menus plus dishes inferred from reviews and photo captions), deduplicates, and filters to dishes with at least one photo so retrieval consistently meets low-latency targets; a fully on-device recognition + matching path (no images sent to the server) with prefetch-on-business-page-open so the scanner has all data locally before it launches; a server-side feature-eligibility gate (GPS proximity to the restaurant, content-availability threshold, frequency capping) that keeps the client thin and lets targeting change without an app release; and a hard-won three-phase matching algorithm (exact → bidirectional substring → Jaro-Winkler similarity) plus an LLM-based menu-normalization pipeline to fix the accuracy and coverage gaps the initial exact-match launch exposed.

Key takeaways

  1. Precompute-and-store-in-Cassandra for low-latency dish retrieval. The data path is: combine all menu sources (owner/partner menus + in-house curated dishes inferred from reviews and photo captions) → deduplicate so each dish appears once → filter to only dishes with ≥1 photo → store the final per-business collection in Cassandra "for fast and reliable retrieval." Pre-processing offline and serving from a fast store is what let Yelp "significantly expand business coverage without building a new system from scratch" while consistently meeting latency targets (Source: this source).

  2. Expand coverage by mining photo captions, not just reviews. Yelp historically inferred menus from review text (powering Popular Dishes), but coverage was limited. Adding user-submitted photo captions as a dish signal both widened the dataset and, because the feature only needs dishes that have an associated photo, aligned the new signal source directly with the feature's own filter criterion (Source: this source).

  3. All recognition and matching runs on-device. Text recognition, text cleanup, and dish matching happen on the client using each platform's native ML frameworks — no images are sent to the server. This minimizes latency and removes any dependence on the user's connection speed at scan time (Source: this source).

  4. Prefetch dish data when the business page opens. Rather than fetch at scanner-launch time, Menu Vision prefetches and locally stores the business's dish collection the moment the user opens the business page, so "when users launch the scanner, all necessary information is already on their device, eliminating the need for additional client-server communications" (Source: this source).

  5. Server-side eligibility gate keeps the client thin. For the contextual "Scan the Menu" prompt at the business entry point, the backend evaluates three conditions per page view: location verification (GPS distance between user and restaurant — only shown to users physically near the restaurant), content availability (query the datastore to confirm the business meets a minimal dish-photo threshold), and frequency capping (limits + cooldown after dismissal). Keeping this logic server-side yields a lightweight client and lets targeting rules change "without needing to release new app updates." Prompting is powered by Yelp's existing educator framework (impression tracking, frequency capping, experimentation) (Source: this source).

  6. Exact-match was too strict; a three-phase matcher fixed it. The initial launch used exact character-for-character matching, which missed real variations ("Garlic Noodle" vs "Garlic Noodles", "Garlic Noodles w/ Pork"). The enhanced matcher runs three phases: Phase 1 — exact match against the primary dish name and all synonyms; Phase 2 — bidirectional substring match (does recognized text contain the dish name, or the dish name contain the recognized text — handles modifiers like "Spicy Pad Thai (V)" → "Pad Thai"); Phase 3 — Jaro-Winkler similarity as a final fallback, chosen because it prioritizes matches at the start of strings, which suits dish names whose key identifiers come first, and it absorbs spelling variants, OCR errors, and transliterations (Source: this source).

  7. LLM pipeline builds a unified, standardized menu per business. To fix coverage/quality gaps, Yelp added a three-step LLM pipeline: standardize partner menus (normalize dish names, extract synonyms, tag price, portion size, dietary labels, calories); process customer language (apply the same workflow to reviews and photo captions to capture how diners actually refer to dishes — "pork ramen" ≈ "Tonkotsu Ramen with Chashu Pork"); intelligent combination (merge and deduplicate both sources, add popularity indicators) into one comprehensive menu per business (Source: this source).

  8. Graceful fallbacks: inventory fallback + QR detection. If the scanner detects no dishes, it displays the restaurant's full dish inventory (popular dishes first) so the feature is still useful when recognition struggles. On iOS it also detects menu QR codes and prompts the user to open the restaurant's digital menu (Source: this source).

  9. Guidance UI improves the input signal. Menu Vision integrates Apple VisionKit's built-in visual guidance ("Slow down", adjust angle) to coach the user into better scanning technique — improving recognition accuracy at the source rather than only compensating downstream (Source: this source).

  10. Staged, monitored rollout. Launched end of October 2025 on both Android and iOS; rolled out gradually stage-by-stage while monitoring client-side logs, server-side load, and error rates. No major bugs or crashes reported; the feature increased retention for users who used it (Source: this source).

Systems extracted

  • Yelp Menu Vision — the end-to-end camera-to-AR dish-discovery feature: on-device OCR + matching, Cassandra dish store, server-side eligibility gating.
  • Apache Cassandra — the fast-read datastore holding the precomputed, deduplicated, photo-filtered per-business dish collections.
  • Apple VisionKit — iOS text-recognition + visual scanning-guidance framework; also used for QR detection.
  • Android CameraX + a third-party ML Kit — the hackathon prototype's camera integration and on-device real-time text recognition on Android.

Concepts extracted

  • On-device ML inference — recognition and matching run on the phone; no images leave the device.
  • On-device text recognition (OCR) — extracting menu text from live camera frames locally.
  • Jaro-Winkler similarity — the prefix-weighted string-similarity metric used as the Phase-3 fuzzy fallback.
  • Multi-phase fuzzy matching — cascading exact → substring → similarity to trade precision for recall in ordered stages.
  • Prefetching — loading dish data at business-page-open so the scanner is instant and offline-capable.
  • Server-side feature eligibility — the backend decides whether/when to surface a feature.
  • Latency budget — the low-latency target the precompute + prefetch + on-device design is built to hit.

Patterns extracted

  • Dedup-filter-precompute to a fast store — combine sources → dedup → filter → store in Cassandra for low-latency reads.
  • Prefetch-on-page-open — fetch the child feature's data when the parent page loads.
  • Server-side feature-eligibility gate — geofence + content threshold + frequency cap evaluated server-side per page view.
  • Tiered exact→substring→fuzzy matching — the three-phase matcher.
  • LLM normalization pipeline for heterogeneous catalog sources — standardize + extract synonyms + merge.
  • Inventory fallback on empty recognition — show full inventory when detection returns nothing.

Operational numbers / specifics

  • Hackathon prototype built in 2 days; Android, using CameraX + a third-party ML kit; tappable pill showed photo count, e.g. "CLAM CHOWDER (969)".
  • Production development began September 2025, target six weeks to ship both iOS and Android at "optimum backend latency".
  • Launched end of October 2025 on Android + iOS; gradual staged rollout.
  • Filter criterion: only dishes with ≥1 associated photo are included.
  • Data sources: owner/partner menus + in-house curated dishes inferred from reviews and photo captions.
  • Matching phases: (1) exact vs primary name + synonyms, (2) bidirectional substring, (3) Jaro-Winkler similarity fallback.
  • Enhanced visual experience shipped in an April 2026 update: dish cards (image, name, photo/review counts, "Popular" badge, price) in a swipeable paginated carousel over the live camera view.

Caveats

  • Tier-3 source and a product-feature story; the load-bearing system-design content is the Cassandra precompute pipeline, the on-device + prefetch latency architecture, the server-side eligibility gate, and the three-phase matcher. Much of the rest is UX/product narrative.
  • No hard latency numbers, dish/business counts, QPS, or Cassandra schema details are disclosed — "low latency targets" and "millions of dishes / restaurants" are stated qualitatively.
  • The Android "third-party ML kit" is not named precisely in the body (consistent with Google's ML Kit text recognition, but treat as unconfirmed); Apple VisionKit is named explicitly.
  • The LLM pipeline's models, prompts, and cadence are not specified.

Source

Last updated · 766 distilled / 2,225 read