Tl;dr: Search ads systems rank and price using only behavioral history: what got clicked before, not what a query or an item actually is. This costs real money: click models underprice exact matches on branded queries by more than 18 percent and overprice substitutes by more than 26 percent, on average. LLMs already know what queries and items mean, and four cached signals (query type, ESCI class, item embedding, user embedding) fix this without running a model at auction time. Amazon, Meta, Instacart, and DoorDash have each built a version of this and published the gains. Geminus and Understudy Labs build it inside your infrastructure.


Ask a person to judge whether a search result is good and they do it instantly. A query for “running headphones” needs earbuds that stay in during movement and handle sweat. A search for “Tide Pods” from someone whose cart already has a fabric softener is probably looking for a specific scent or size. A user whose history shows a treadmill, running shoes, and a protein powder shaker is obviously a fitness household. None of this requires seeing what they clicked before. It requires knowing what things are.

Search ads systems don’t know what things are. They know what users clicked before.

A ranking model built on behavioral data can notice that certain words and certain purchases co-occur, and so can a relevance model, whose job is to decide whether an item is a good enough match to be eligible for a given auction at all. Neither knows why. A relevance model doesn’t know that running headphones need to stay put during movement, only that people who search the phrase tend to buy certain items. On a popular query with years of click history, that doesn’t matter: enough clicks have piled up to fake it. On a new listing with zero history, or a reasonable phrasing nobody has tried before, it matters completely. A model with real world knowledge judges the item correctly on the first try, because it understands the requirement. A behavioral model just waits, hoping enough people click before it has to guess.

The same gap shows up with items. A phone case and a phone are complements. Two different-brand running shoes are substitutes because they serve the same purpose. A brand sits at the premium end of its category or the budget end. None of this comes from watching what people happened to buy together on one platform. It comes from understanding what these products actually are: a kind of knowledge absorbed from a vastly larger body of product descriptions, comparisons, and reviews than any single marketplace generates on its own.

And it shows up with users. A household whose purchase history shows a treadmill and running shoes is, to anyone who knows what these things are, obviously someone to show running gear and sports bottles to, even if they have never once searched or bought in these categories before. A behavioral model has no such shortcut. It can only guess by looking at what similar users did: a noisier, slower proxy for something a person, or an LLM, understands immediately.


In search, the query is the most direct signal of what a user wants at this moment. History matters, and it shapes which ad wins among eligible candidates. But which products get retrieved and which are allowed to compete in the auction depends entirely on the query and its relationship to each product. A system that cannot tell whether “Sony WH-1000XM5” is a branded specific search or whether Bose QC45 is a substitute for it cannot make correct retrieval or eligibility decisions, no matter how much it knows about the user making the request. And the same is true further down the pipeline: a ranking model that doesn’t understand what the user is looking for right now, or whether a given product genuinely answers it, cannot predict click or conversion probability accurately, no matter how well it knows the user. Getting the query and product relationship right is the prerequisite for all the others.


What this costs

The gap between world knowledge and behavioral data shows up in three concrete places.

1. Retrieval can’t expand past what’s already there. In most retail media networks (RMNs), and in essentially all third-party retail media platforms (RMPs), ad retrieval is built on top of organic retrieval. Third-party RMPs have no access to real-time retrieval at all: they receive the candidate set organic already produced and work within it. Even RMNs that run their own ad infrastructure typically use organic retrieval as the starting point. Organic retrieval wasn’t built to maximize advertiser coverage. It’s built to return what best answers the query. Substitutes and complements in adjacent categories (e.g., a competing running shoe brand on a running shoe query), items users would genuinely consider, are often never retrieved. An advertiser with a willing bid, a relevant item, and an active campaign never enters the auction, because the retrieval system wasn’t looking for them. Without a reliable way to judge whether an item from an adjacent category genuinely matches a query, expanding retrieval risks surfacing irrelevant results. So the safe choice has been not to try. An entire inventory of potentially eligible ads sits permanently invisible, regardless of how much advertisers would pay to show them.

2. Eligibility rules can’t make fine distinctions. A common assumption is that pCTR handles relevance on its own: irrelevant items get low click rates, lose auctions, and disappear. In practice, this fails in several ways. Users click irrelevant ads out of curiosity or confusion, especially in position one, and the model learns from the clicks that the item performs well. A high bid can push a low-relevance item into a top slot regardless of its predicted click rate. On queries with thin coverage, a weakly relevant product can win simply because nothing else is competing. GMV suffers, session quality drops, and users gradually stop trusting the page. Relevance-based eligibility is a guard against all of this. It’s not there to influence ranking. It exists to protect user experience, GMV, and marketplace health.

However, even for items organic does retrieve, the rules that guard eligibility can’t encode the distinctions that matter. The same rule applies to a narrow branded query and a broad exploratory one, even though each calls for a different answer. A substitute brand is a perfectly good result on a broad discovery search and a bad one in position one on a specific branded query. More sophisticated systems use a relevance model, trained on human or LLM-generated labels, that scores each query-item pair and gates eligibility with a threshold. That’s better than a simple category rule. However, the underlying problem remains. A relevance score of 0.85 might mean 70 percent chance of exact match, 20 percent chance of substitute, a few percent irrelevant, all compressed into one number. Setting a threshold on the number is always a compromise. Set it high enough to filter out the irrelevant items and good substitutes sitting just below the cutoff get excluded along with them. Set it lower and irrelevant items start appearing in positions they have no business occupying. There’s no setting that does both well, and the threshold drifts out of calibration every time the underlying model is retrained.

3. Ranking and pricing models work with signals that don’t encode what they need. Even in auctions that do run, with the right candidate set and sensible eligibility gates, two models sit at the center of every auction decision. The pCTR model predicts how likely a user is to click a given ad. Multiplied by the advertiser’s bid, it determines which items win which positions: the highest score wins the top slot. The pCVR model predicts how likely a click is to convert into a purchase, and advertisers use it to calibrate how much they should bid in the first place to hit their return targets. Both are working with signals that don’t tell them what they most need to know: whether this is a branded, specific search or a broad exploratory one, whether the item is an exact match, a substitute, a complement, or something unrelated. The models average over all of these situations because no explicit signal distinguishes them. The result is documented in measured production data. Click prediction models underprice exact matches on branded narrow queries by more than 18 percent on average. They overprice substitutes on broad queries by more than 26 percent. Those are averages, and averages hide what is happening at the individual item level, where a meaningful share of rankings are fully inverted. The wrong item wins position one. The right item sits below it. The model then trains on the clicks produced by the wrong ranking and concludes the weak match was good, when it only performed well because it was shown first. Every retraining cycle inherits the same distortion, a little worse.

All three problems come from the same missing piece. The systems do not know what queries mean, what items are, how well they match, or what a user’s history implies about them. They know only what happened before on this one platform.


What LLMs change

Large language models know what things are. They were trained on a vastly larger body of text than any single marketplace generates, e.g., product descriptions, reviews, category taxonomies, brand comparisons, consumer guides. An LLM knows why running headphones need to stay put. It knows that Persil and Tide are substitutes. It knows that a camping stove query implies interest in fuel canisters. It knows that a household with treadmill and protein powder purchases has fitness intent, without needing to see what other similar households bought.

This knowledge can be applied offline, to the full catalog and the full query space, and the results cached as discrete labels. The first step is query type classification: for each head and torso query, an LLM determines whether it is a branded search or a non-branded one, and whether it is narrow and specific or broad and exploratory. This single label changes almost every downstream decision. A branded narrow query for a specific product calls for exact matches in top positions. A broad non-branded query is where substitutes and complements belong. Without this classification done explicitly and cached, every system downstream has to guess at the distinction from indirect signals, and gets it wrong often enough to move bid prices by double digits. The second step is query-item classification: for each pair of query and item across the catalog, an LLM assigns an ESCI label. For the queries and query-item pairs that account for roughly 80 percent of traffic, both steps can be done in advance. When a request arrives, using the results is a lookup. No model runs at serving time.

Both sets of labels are discrete: query type is branded or not, narrow or broad; item relevance is Exact match, Substitute, Complement, Irrelevant, or Offensive (i.e., a brand safety violation). Each pair gets one of these. A continuous score would compress all four cases into one number, and the threshold you set on the score is always a compromise: calibrated for one situation, it miscalibrates another. A discrete label supports a rule that says exactly what it means. A rule that says “top position, branded query, exact match only” does exactly what it says, with no approximation and no accumulated error rate between retrains.

The third step is the item-semantic embedding: each item in the catalog is run through a language model once, producing a vector that captures what the item actually is based on its name, brand, category, and price. Two running shoes from different brands end up with similar vectors because they serve the same purpose, even if no user on this platform has ever bought both. This representation captures meaning no amount of behavioral co-occurrence data can produce, because it comes from understanding what the item is rather than from watching what people clicked alongside it.

The fourth step is the user embedding: a compact representation of what a user’s interaction history implies about them, built from the same item-semantic vectors. A language model reads the sequence of items a user has engaged with and produces a summary of their interests grounded in world knowledge. A user whose history shows a treadmill, running shoes, and a protein powder shaker is, to the model, a fitness household. That inference is available on day one, before the platform has accumulated any behavioral data connecting this specific user to any specific product category.

Using frontier language models to classify 80 percent of queries against the full advertised catalog is prohibitively expensive at production scale: e.g., a large retail platform might have tens of millions of query-item pairs to label across head and torso traffic alone, and frontier inference costs make this infeasible as a recurring operation.

This is where Understudy Labs comes in. Understudy runs a full optimization ladder: a frontier model panel establishes ground truth on a sample, checking each other’s labels and treating agreement as the verified standard; open-weight candidate models are screened against the standard on quality, cost, and latency; prompt optimization via GEPA, a gradient-free reflective optimizer, closes much of the remaining gap with no weight updates at all; and supervised fine-tuning handles the residual error slices where prompt optimization is not enough. In a validation study on a real marketplace specificity classifier, GEPA alone moved balanced accuracy from 76.7 percent to 82.7 percent and macro-F1 from 72.4 percent to 81.2 percent on held-out data, before any fine-tuning. The result is a compact open-weight model the retailer owns, running inside their own infrastructure at roughly 15 percent of frontier inference cost. Because everything runs locally, query and item data never leaves the platform’s systems, and there is no risk of the data being incorporated into the frontier models that set the original standard. A companion post covers the Understudy architecture and the validation results in detail. Amazon, Meta, Instacart, and DoorDash have each independently built the same architecture for the same reason. Amazon reported a 60 percent improvement in search relevance and a 0.7 percent sales uplift on deployed traffic. Instacart cut search complaints in half and expanded query understanding coverage from 50 percent to over 95 percent. Meta measured a 5 percent lift in ad conversions on Instagram in the quarter their foundation model shipped. DoorDash has validated both halves of this argument in separate production papers: LLM-derived query intent classification (deployed on 95 percent of search traffic, with an ablation showing +8.3 percentage points from catalog grounding alone) and LLM-generated ordinal relevance labels for over 100 million query-item pairs, folded directly into the ranking value function alongside CTR and conversion signals. These gains came on top of systems that already had ads-specific retrieval, query tagging, sophisticated relevance models, strong query-item and user-item embeddings, and mature eligibility and auction infrastructure. They were not starting from organic retrieval and category rules. A platform that is still at this starting point therefore has considerably more room to move.

One honest counter-result belongs in this picture. LinkedIn’s CADET, a pure behavioral sequence model with no semantic signal anywhere in its inputs, achieved an 11 percent CTR lift over their prior production system. That’s real evidence that a well-executed collaborative sequence model can perform well on its own. But the “semantic IDs” CADET tested were quantized from behavioral and collaborative embeddings, compressing learned interaction patterns into discrete codes. They weren’t drawn from outside world knowledge about what a brand or category means. CADET is a negative result for one specific way of adding discreteness, and a strong result for collaborative modeling. It’s not a test of the query-type and ESCI-class argument this post is making.


Why this was impossible before, and what changed

This problem persisted for years even though its costs were visible. There’s a reason: fixing it required expanding ad retrieval and eligibility beyond what organic search already provided, and the expansion ran directly into an organizational wall.

In every retail media network, the organic search team owns user experience. They guard it for good reason. An ads team that expands retrieval into adjacent categories, or applies a relevance model to decide which items are eligible, is introducing a new source of bad ads: items organic would never show, admitted by a model that is imperfect by construction. When these items degrade search quality, the organic team escalates, and the ads team is held responsible for the model’s errors. The rational response has been to stay inside organic’s safe perimeter and not try. Nobody wanted to own this risk.

For third-party RMPs, the problem is worse. They have no relationship with the retailer’s organic team at all. They can’t negotiate what gets admitted. They can’t be held accountable for retrieval they don’t control. The only safe option has been to take whatever organic surfaces and work within it.

A relevance model doesn’t solve this because it can’t make a credible promise. A model score of 0.85 might be right most of the time. However, neither the ads team nor the organic team knows which specific items it got wrong. There is no auditable answer to “why did this item appear here.” The organic team can’t agree to something they can’t inspect, and the ads team can’t commit to outcomes they can’t predict.

Explicit, cached, LLM-derived labels dissolve this barrier. When eligibility runs on a discrete ESCI class and a query type tag, both produced offline and verifiable, the rules that govern which items appear in which positions can be written down, reviewed, and agreed on before a single ad is served under the new system. The organic team can see exactly what the rules say. The ads team is accountable for following the rules. Their accountability is no longer contingent on model behavior they can’t predict. Adjusting the tradeoff between ad revenue and user experience becomes a policy decision both teams can negotiate rather than a retraining cycle they can only guess at. This is what makes expansion past organic’s perimeter possible for the first time.


How it changes each layer

1. Retrieval. Query type tags and ESCI labels make it possible to expand the candidate set deliberately. Items in adjacent categories can be retrieved and judged against the query with a known relevance class, rather than excluded by default. An advertiser with a relevant substitute or complement and a willing bid can enter auctions they could never reach before. And for head and torso queries, the cache collapses retrieval and eligibility into a single lookup. There’s no separate retrieval step followed by a separate eligibility check: the cache already contains, for each query, exactly which items are eligible for which positions. Third-party RMPs benefit from this most directly. They currently have no retrieval system of their own. With this architecture, they gain one, in the form of a cache that was built with world knowledge and agreed on with the organic team before any ad was served under it.

2. Eligibility. The ESCI class per query-item pair, combined with the query type tag, supports eligibility rules that are explicit and position-aware. Which items can appear in which slots, on which kinds of queries, is a policy that can be stated, agreed on in advance with the organic search team, and audited at any point. The rules cube (i.e., the matrix of query type by relevance class by slot position) is tunable. If marketplace health metrics show a tradeoff going the wrong way, tighten the relevant rule. If demand signals suggest more inventory should compete in a category, loosen it. Either adjustment is immediate and has a known, observable effect, with no threshold drift, no recalibration cycle, and no invisible error rate accumulating between model retrains.

3. pCTR and pCVR. All four LLM-derived inputs, query type, ESCI class, item-semantic embedding, and user embedding, feed into click prediction and conversion prediction models as explicit features. Query type and ESCI class eliminate the averaging that produces the measured bias: the model now knows whether this is a branded narrow query or a broad exploratory one, and whether the item is an exact match or a substitute, rather than collapsing both distinctions into one blended signal. Item-semantic embeddings let the model judge relevance between query and item from meaning rather than from patterns of co-occurrence, which matters most for new listings and unusual phrasings where behavioral history is thin. User embeddings let the model personalize from what a user’s history implies rather than only from what users statistically similar to them happened to do. The pCTR model stops putting the wrong item in position one. The pCVR model stops systematically driving overbids on weak matches and underbids on exact ones, which means campaigns hit their return targets instead of requiring constant manual correction. The models stop inheriting distorted training signal from their own wrong rankings. The loop breaks.

Adding these four signals to the existing ranking model is the first and most valuable step, and it doesn’t require building any new model architecture. For platforms that want to go further, a teacher/student architecture takes these same signals as inputs, fine-tunes a teacher model on real click and conversion outcomes to establish an accuracy ceiling, and then distills a fast two-tower student from the teacher’s predictions, producing cached item and user vectors that serve at auction latency. A companion post covers this architecture in detail.


What Geminus and Understudy Labs offer

We build and deploy this infrastructure inside your infrastructure.

1. Query and query-item classification. We produce ESCI labels for your catalog and query type tags for your head and torso query volume, roughly 80 percent of traffic. Labels come from Understudy’s LLM consensus pipeline, cached, and served as zero-latency lookups. Nothing leaves your systems. No data retention.

2. Eligibility rules. We design the query type, ESCI class, and slot position rule matrix with you, agreed with your organic search team before any ad is served under the new system. The rules are deterministic and auditable. Changing the tradeoff between user experience and revenue means changing the rule, which takes effect immediately and is observable before and after.

3. Ads-specific retrieval. For platforms ready to expand beyond what organic surfaces, we scope and label adjacent-category items, the inventory that currently never enters an auction regardless of advertiser bid.

4. pCTR and pCVR feature improvement. We produce the LLM-derived features your ranking and bidding models currently lack: query type, ESCI class, and their interaction. These are the features that remove the averaging producing documented bid and ranking bias.

5. Tail query handling. The roughly 20 percent of traffic outside the head and torso cache is handled in two ways. Query mapping finds semantically similar cached head or torso queries and applies their labels directly, with a confidence threshold that gates which mappings are trusted. For queries where no confident mapping exists, a compact query-item relevance model trained on the labeled head and torso set produces a relevance score used with conservative thresholds, admitting fewer items and only the clearest matches, so the risk of a bad ad in a low-coverage situation stays bounded. Organic retrieval handles the remainder as it does today.


How to get started

If you operate your own search and ads infrastructure, we scope the work, run the Understudy labeling pipeline inside your infrastructure, design the eligibility rules jointly with your organic team, and run a bounded pilot on a single query segment before broader deployment. From scoping call to pilot results is about ten weeks.

If you work with a third-party retail media platform, your path runs through your RMP. We can present this directly to your RMP as a platform capability for their full client base.

If you are a retail media platform, this is a product differentiator you deploy across clients. We build the labeling pipeline and eligibility rule engine as a parameterized platform feature. The per-client work is data access and organic team negotiation. The platform investment amortizes across your entire book.

To start a conversation, share your total annual search ad revenue and, if available, fill rate by query segment and CPC distribution on branded queries. We produce a platform-specific estimate within one week.

vadim@geminusdata.ai