Skip to content

SYSTEM Cited by 1 source

Pinterest User Interest Clusters (UIC)

Definition

User Interest Clusters (UIC) is Pinterest's stateful, per-user interest representation at the core of the Pinner Progression program. Instead of representing a user with a single global interest vector, UIC represents them as a set of interest clusters, each corresponding to a distinct use-case (DIY inspiration, recipe discovery, apartment decorating, marathon training, …). Each cluster carries lifecycle metadata so the recommendation stack can reason about where an interest sits in its lifecycle — newly discovered, established habit, or decaying (Source: sources/2026-07-27-pinterest-pinner-progression-better-use-case-representation-driving-weekly-active-user-growth).

UIC is the concrete implementation of two wiki concepts introduced by the same source: use-case representation and interest-lifecycle modeling.

Lineage: PinnerSage → OmniSage → UIC

UIC builds on Pinterest's history of multi-modal user representation:

  • PinnerSage introduced representing each user with multiple embeddings by clustering their engagement history and using cluster medoids as retrieval queries — a step beyond single-embedding models.
  • OmniSage advanced this via multi-entity graph representation fusing visual/semantic features, interaction-graph signals, and Pin-Board topology into a unified embedding whose "closeness" encodes functional utility, not just visual similarity.
  • UIC innovates on top in three ways (below).

Three innovations over PinnerSage/OmniSage

  1. Personalized clustering over engaged content only. Clusters are built over only the Pins a user engaged with, not the global content catalog. This makes the problem tractable and clusters semantically coherent — each cluster is a use-case the user is actively pursuing. The same Pin (same OmniSage embedding) may land in different clusters for different users, capturing the personal nature of "use-case."
  2. Dynamic cluster count. No fixed k. The clustering algorithm picks the natural number of clusters per user from a coherence threshold — a new user might have ~2 clusters, a power user ~15.
  3. Stateful lifecycle metadata. Each UIC carries temporal / behavioral metadata (recency, frequency of engagement) so downstream layers can treat interests differently by maturity. The post is candid this is a starting signal, not a full solution to distinguishing fleeting curiosity from an emerging habit.

Signal construction

Each UIC is defined by a medoid plus a group of landmark Pins, all in the OmniSage embedding space. Construction:

  1. Collect the user's recent engagement sequence — last ~500 actions (closeups, saves, clicks). This is a consumer of the user-sequence platform substrate.
  2. Run complete-linkage agglomerative hierarchical clustering on the action embeddings: start with every engaged Pin as a singleton, greedily merge the two most similar clusters until no remaining pair exceeds similarity threshold τ (or the cluster count hits an upper bound).
  3. Complete linkage defines similarity between two clusters as the similarity of their least similar pair — a merge is allowed only if every point in cluster A is sufficiently similar to every point in cluster B, using cosine similarity between OmniSage vectors. This yields tight, coherent clusters.

A worked example in the post: five engaged Pins (three cats, two jeans) with cat-cat and jeans-jeans cosine similarities of 0.7+ but cross-category below τ merges into exactly two clusters {cats}, {jeans} — the final {cats}-vs-{jeans} complete-link similarity falls below τ so the algorithm stops.

System-level integration

UIC is used as a shared abstraction across the whole Home Feed serving funnel (retrieval → ranking funnel). A key design decision was to externalize UICs to a shared feature store rather than couple them to a single model — see externalize-signal-to-shared-feature-store. UIC features are fetched once (from GSS) and attached to the User UFR node, so every stage can read them without redundant fetch latency.

Stage How UIC is used
Retrieval (CLR) UIC clusters replace the static "followed-interests" condition; controls candidates allocated per use-case (e.g. sample 5 of 10 clusters); frontier sampling for exploration; scope to active clusters to cut overfetch.
L1 Utility Control layer between LWS and ranking; annotates each candidate with its UIC (cosine to medoids) and applies a penalty-based diversity discount by UIC_dupes.
Ranking (Pinnability) State-dependent utility weights — curiosity signals weighted higher for nascent interests, commitment signals for mature ones — without adding prediction heads.
Diversity (SSD blending) UIC-aware penalty; each Pin assigned its best-matching medoid above 0.85 similarity; penalize proportional to UIC_coverage already selected.

Impact

Qualitative per Pinterest policy: UIC-conditioned retrieval cut overfetch and infrastructure cost while keeping candidate volume unchanged; UIC-aware SSD produced meaningful engagement gains, much more diversity of interacted content, and more longer sessions. The overarching motivation is retention, not immediate engagement — see retention-as-first-class-objective.

Seen in

Last updated · 766 distilled / 2,225 read