Pinner Progression: Better Use-Case Representation Driving Weekly Active User Growth at Pinterest¶
Summary¶
Pinterest introduces Pinner Progression, a program that reframes the Home Feed recommendation system around retention as a first-class objective rather than pure engagement. Its core signal is the User Interest Cluster (UIC) — a stateful, per-user set of interest clusters built by running complete-linkage agglomerative hierarchical clustering over each user's last ~500 engaged Pins in OmniSage embedding space. Each UIC carries temporal/behavioral lifecycle metadata (recency, frequency) so downstream layers can reason about whether an interest is nascent, an established habit, or decaying. UIC is externalized to a shared feature store, attached once to the user node, and consumed as a common abstraction across every stage of the serving funnel: Conditional Learned Retrieval, the L1 Utility control layer, ranking, and the SSD blending/diversity layer. This is Part 1 of 2; Part 2 will cover predicting unseen UICs and systematic interest exploration.
Key Takeaways¶
-
Engagement ≠ retention — optimizing the feed purely for immediate engagement signals (clicks, saves, downloads, closeups) feeds users more of what they already like but never helps them find something new for next time; a user can save ten sourdough recipes today and churn next month anyway (Source: Section "When Engagement Optimization Is Not Enough").
-
Use-case adoption is a durable proxy for user value — the durable, sustained actions (multiple Repins, long clicks) that mark use-case adoption predict retention. The relationship is non-linear: modest for moderate increases in adoption, then accelerating sharply at the top decile, suggesting a threshold effect where breadth of adoption compounds retention (Source: Section "When Engagement Optimization Is Not Enough").
-
Point-estimate models miss the lifecycle of interests — existing systems (TransAct-style real-time action sequences, learned two-tower retrieval, diversification) model the user as a set of recent actions. They cannot distinguish an accelerating "apartment decorating" era (just signed a lease) from a decaying "sourdough" phase (mastered months ago), even though only the former sustains long-run visitation (Source: Section "When Engagement Optimization Is Not Enough").
-
UIC = personalized clustering over engaged content only — rather than clustering in the global embedding space, Pinterest clusters only the Pins a user engaged with, making the problem tractable and clusters semantically coherent per-user. The same Pin with the same OmniSage embedding can land in different clusters for different users, capturing the personal nature of a "use-case" (Source: Section "From PinnerSage/OmniSage to UICs").
-
Dynamic cluster count — users don't have a fixed number of interests (a new user might have two; a power user planning a wedding, remodeling a kitchen, and marathon training might have fifteen). The clustering algorithm determines the natural number of clusters per user via a coherence threshold rather than a fixed k (Source: Section "From PinnerSage/OmniSage to UICs").
-
Stateful lifecycle metadata — each UIC carries temporal/behavioral metadata (recency and frequency of engagement) that gives downstream layers a structured signal to treat interests differently by maturity. The post is candid that this does not fully solve distinguishing fleeting curiosity from an emerging habit — it provides a starting point (Source: Section "From PinnerSage/OmniSage to UICs").
-
Complete-linkage agglomerative clustering — start with every engaged Pin as a singleton cluster, greedily merge the two most similar clusters until no remaining pair exceeds a similarity threshold τ (or cluster count hits an upper bound). Complete linkage defines cluster similarity as the similarity of the least similar cross-pair, using cosine similarity between OmniSage vectors — a merge is allowed only if every point in A is sufficiently similar to every point in B (Source: Section "Signal Construction").
-
OmniSage "closeness" encodes functional utility — a hiking boot and a trail-mix bar are neighbors because they are co-curated on "Hiking Trip" boards and co-engaged by the same users. Clustering in this space produces clusters corresponding to coherent life projects (planning a camping trip, redecorating a bedroom) rather than visual categories (Source: Section "The Embedding Space").
-
Externalize the signal to a shared feature store — one of the most important design decisions was to externalize UICs to a shared feature store attached to the User UFR node rather than coupling them to a single model. Because UIC is fetched once, candidate annotation happens once, avoiding redundant fetch latency across Home Feed stages (Source: Section "System-Level Integration"; note "we fetch from GSS once").
-
UIC-conditioned retrieval + budgeting — CLR replaced the static "followed-interests" condition (which skewed to dominant interests and didn't evolve) with UIC clusters, controlling how many candidates are allocated per use-case (e.g. sample 5 of a user's 10 clusters). Scoping retrieval to currently active clusters (recent engagement, above coherence threshold) reduced overfetch — broadly relevant but unused candidates — yielding meaningful infrastructure cost savings while keeping total candidate volume to downstream stages unchanged (Source: Section "Retrieval").
-
L1 Utility as an active control surface — the control layer between lightweight scoring (LWS) and full ranking annotates each candidate with its UIC (via cosine similarity to UIC medoids) and applies a penalty-based diversity mechanism: Pins duplicating an already-well-represented UIC get a discounted score, so under-represented interests survive to the ranker.
L1_scoreis discounted by a function ofUIC_dupes(how many higher-scoring Pins share the same UIC) (Source: Section "L1 Utility"). -
State-dependent ranking utility weights — ranking converts raw Pinnability predictions into final scores using UIC lifecycle state: for a nascent interest, curiosity signals (clicks, closeups) matter more than commitment signals (saves); for a mature interest, the reverse. Making utility weights state-dependent aligns the objective with user needs without adding new prediction heads to the ranking model (which would raise serving latency) (Source: Section "Ranking").
-
UIC-aware SSD diversity — Sliding Spectrum Decomposition previously optimized Pin-level variety (adjacent Pins look different) but had no notion of higher-level use-cases, so a feed could look visually diverse yet be dominated by one interest. A UIC-aware penalty (assign each Pin its best-matching cluster medoid above a 0.85 similarity threshold; unmatched Pins fall into a default group) penalizes Pins proportionally to how much their UIC is already represented in the selected set — so newer interests surface naturally in later feed positions once dominant clusters accumulate penalty (Source: Section "Diversity").
-
Online results — UIC-aware SSD produced meaningful engagement gains, much more diversity of interacted content, and increases in longer sessions, confirming that balancing use-case representation translates to deeper engagement (Source: Section "Diversity").
Systems / Concepts / Patterns extracted¶
- Systems: User Interest Clusters (UIC), OmniSage, PinnerSage, Conditional Learned Retrieval (CLR), Sliding Spectrum Decomposition (SSD).
- Concepts: use-case representation, interest-lifecycle modeling, retention as a first-class objective, complete-linkage agglomerative clustering; extends retrieval → ranking funnel, two-tower architecture, feature store, user event sequence.
- Patterns: externalize signal to a shared feature store, state-dependent ranking utility weights, penalty-based diversity in the blending layer, retrieval budgeting to cut overfetch, frontier sampling for exploration.
Operational Numbers¶
| Metric | Value |
|---|---|
| Engaged-history window per user | last ~500 actions (closeups, saves, clicks) |
| Cluster count | dynamic per user (coherence threshold), example range ~2 to ~15 |
| Clustering algorithm | complete-linkage agglomerative hierarchical clustering |
| Distance metric | cosine similarity in OmniSage embedding space |
| SSD UIC-assignment similarity threshold | 0.85 (best-matching medoid; else default group) |
| Retrieval budgeting example | sample 5 of a user's 10 clusters |
| Feed pipeline stages touched | retrieval (CLR) → LWS → L1 Utility → ranking (Pinnability) → SSD blending |
| Feature-store fetch | fetched once (from GSS), attached to User UFR node |
Caveats¶
- Numbers are largely qualitative per Pinterest's disclosure policy ("meaningful gains", "significant cost savings"); no absolute percentages given for online lifts.
- The post is explicit that lifecycle metadata is a starting signal, not a full solution to curiosity-vs-habit disambiguation.
- This is Part 1 of 2 — prediction of unseen UICs (geometric prediction, model-based serendipity, LLM interest reasoning) and an RL exploration feedback loop are deferred to Part 2.
Source¶
- Original: https://medium.com/pinterest-engineering/pinner-progression-better-use-case-representation-driving-weekly-active-user-growth-at-pinterest-bd2131ab238a?source=rss----4c5a5f6279b6---4
- Raw markdown:
raw/pinterest/2026-07-27-pinner-progression-better-use-case-representation-driving-we-7b35b642.md
Related¶
- systems/pinterest-user-interest-clusters
- systems/pinterest-omnisage
- systems/pinterest-pinnersage
- systems/pinterest-conditional-learned-retrieval
- systems/pinterest-sliding-spectrum-decomposition
- use-case-representation
- interest-lifecycle-modeling
- retention-as-first-class-objective
- complete-linkage-agglomerative-clustering
- companies/pinterest