Shifting Marketplace Discovery from Engagement to Semantic Understanding
Traditional systems frequently treated discovery as an engagement optimization problem, leading to suboptimal results for nuanced queries.
DoorDash moves beyond engagement metrics, leveraging LLMs to understand shopper intent and item meaning for a diverse marketplace.
Traditional systems frequently treated discovery as an engagement optimization problem, leading to suboptimal results for nuanced queries.
To capture every shoppable moment, the company now facilitates grocery, retail, pets, and gifting services alongside its original restaurant offerings.
Consumers frequently engage in broad shopping missions that span multiple sessions or days, where a single query like 'new puppy' can involve various product categories and collections.
Supervision, catalog semantics, semantic personalization, and steerable content generation serve as the foundation for the firm's mapping of the customer journey.
Supervision involves teaching systems what constitutes 'good' for ranking, while catalog semantics focuses on understanding the inherent nature of items, and semantic personalization tailors results to specific user intent, with steerable content generation producing dynamic results.
Popular regular spaghetti frequently obscures 'gluten-free pasta' in search results because engagement signals bias rankings toward high-volume items.
Engagement metrics are inherently biased by factors such as exposure, position, and price, underscoring the need for graded relevance to distinguish between exact matches, substitutes, and irrelevant items.
If we only rank from engagement, very popular regular spaghettes might show up and because it sells well, you know, it may continue to show up.
High-quality training labels are generated by leveraging LLMs to perform expensive reasoning offline, which is then distilled into lightweight models to supplement human ground truth.
The process involves performing expensive reasoning offline to reconcile human feedback with category models, creating a golden dataset for distillation into lightweight, quickly serviceable models.
Standard embeddings often collapse the distinction between related and matching items at e-commerce scale, posing a significant challenge for retrieval systems.
DoorDash addresses this by employing a two-stage contrastive learning approach: first, global geometry shaping with multi-level supervised contrastive loss, followed by curriculum training using hard negatives that the model previously confused.
Implementing a two-stage process in the retrieval layer has led to a 2.3% improvement in NDCG for DoorDash.
An ordinal relevance tower, which leverages LLM-graded labels, has been added to rankers, sharing bottom layers to distill semantic fit alongside engagement metrics like click and conversion.
A value function is employed to fine-tune the balance between engagement and relevance across different surfaces, utilizing LLMs for reasoning and understanding rather than replacing the entire retrieval and ranking infrastructure.
Semantic IDs provide a short hierarchical code, analogous to taxonomy, where prefixes define broad neighborhoods and later tokens capture finer distinctions, emerging from unlabeled data to categorize items like hot sauces by specialty (e.g., Mexican, Caribbean, Korean).
This system facilitates cross-category relationships, solves cold start problems, provides coverage for tail items, and enables reverse auditing to compare human labels against semantically learned labels.
What it does is it gives us a you know a short hierarchical code that's analogous to our taxonomy but uh the ability to control the fine grain nature of it.
Semantic IDs have significantly improved DoorDash's ranker, increasing its Mean Reciprocal Rank (MRR) by 4% to 5% and leading to measurable conversion gains.
In query reformulation, the system maps queries to semantic neighborhoods to suggest more relevant items, preventing suggestions for out-of-stock items and resulting in substantial MRR gains for query suggestions.
Get this channel’s next video as an article by email.
This approach builds on well-understood agent-context memory concepts to enrich personalization.
This framework distinguishes between long-term memory, which tracks durable preferences from past orders and searches, real-time context managing session-specific interactions like cart state, and stated preferences capturing direct input from agentic interactions such as Ask DoorDash. Memory blocks are decoupled, allowing new user dimensions to be added without reinterpretation.
These include human-readable text summaries, latent vector embeddings for integration with retrieval and ranking systems, and graph and hierarchical approaches to map complex relationships between brands and taxonomies.
Graph-based embeddings are currently outperforming existing taxonomy-based embeddings in retrieval tasks, demonstrating their effectiveness.
This memory framework supports personalized collections, agentic personalization, and provides crucial features for downstream models and LLM applications.
DoorDash has transitioned from fixed content libraries to generating dynamic, consumer-level collections through offline LLM processing.
This system tailors content to specific shopper affinities, such as plant-based diets or particular household pets, with testing in the pets vertical showing a 1% increase in order rates and a 6% increase in active users, achieved through batch processing that enables sophisticated reasoning while maintaining system performance.
Discovery is fundamentally a semantic understanding problem rather than solely an engagement metric.
The optimal architecture rarely involves direct online LLM calls; instead, reasoning should be distilled offline into scalable primitives like semantic IDs and memory blocks, which enables smaller, faster models to serve these insights consistently across retrieval, ranking, and content generation.
Answers come from the transcript, with the exact spot cited.
Want the next video from AI Engineer too?
When AI Engineer publishes, we'll write it up like the one you just read and email it to you.
AI Engineer published 29 in the last 7 days.
Spotify Transforms Recs with LLM-Native AIAI Engineer4 hours ago · 19:40 · 147 views · Created 3 hours ago
LLM Recommenders Poised to Be AI's Biggest Consumer AppAI Engineer4 hours ago · 18:00 · 167 views · Created 4 hours ago
Moonlake AI's Mission: Embodied IntelligenceAI Engineeryesterday · 51:35 · 3.8K views · Created 16 hours ago