Tl;dr: Every layer of how CPG audiences get built today adds a gap between what a brand needs and what it ends up buying. New-to-brand audiences have to be trained on households that already converted, and by the time you look at them they’re current customers, not people about to convert, so brands settle for proxies. Lookalike expansion doesn’t fix this: it finds people who resemble current customers, not people on their way to becoming them. Statistical models trained on transaction data can’t learn relationships they haven’t counted often enough, e.g., gym memberships and sports nutrition; a language model fine-tuned on real outcomes already knows the connection. Attribution can’t tell what a campaign caused from who would have bought anyway. A randomized holdout can. Geminus closes all four gaps at once: an AI agent translates the objective, a language model fine-tuned on the behavior of converters before they converted scores the full cross-retailer panel, and a holdout proves the lift.


When a brand buys an audience, what it’s actually buying is usually a stand-in for the audience it needs. It gets assembled from whatever signals are available and expanded using whatever data the vendor has. Nobody can tell you afterward whether it worked. Both sides know it, at some level. The transaction happens anyway, because there hasn’t been an alternative.

The transaction-level data needed to find a real audience, including real purchase histories and real household behavior tracked over time, exists. It’s commercially available, and has been for a while. The problem is structural. It runs through every layer of how audiences get built today, from the initial request to the final measurement. Each layer adds a gap between what a brand actually needs and what it ends up buying. There are four of these gaps, and closing each one changes what an audience product can actually do.

1. The proxy problem

A brand wants to grow its buyer base. The most efficient path is not to advertise to all non-buyers, most of whom have no real likelihood of becoming customers anytime soon, but to find as many as possible of the households already somewhere on a trajectory toward a first purchase and concentrate the message and the marketing spend on them. This audience cannot be built by conventional means.

The most direct approach would be to take a list of households that became new-to-brand buyers in the past, understand what they looked like in the months before they converted, and find other households showing the same pattern right now. But this runs into a technical problem that sounds simple and turns out to be decisive: real household-level transaction history is rare, and even where it exists, it doesn’t scale to the size an audience needs. It’s also just hard to build: instead of reusing pre-built, pre-modeled features, like propensities or user-item embeddings, you need new longitudinal features for every new segment you expand into. What’s actually available, for the converters and for the households you’d expand to, is current, static data: aggregated features built from consumer behavior, e.g., purchases, visits to places of interest, browsing activity, and demographics. Customers who became new-to-brand buyers in the past are, by definition, no longer new-to-brand customers. They are current customers. A model trained on them, using their current purchase behavior and current demographics, is not a model of the people on their way to becoming buyers. It’s a model of existing buyers. The signal that predicted conversion, what a household looked like just before it changed, is gone by the time the modeling starts, replaced by the static characteristics of someone who has already made the change.

The same problem shows up in other audience strategies too. An upsell audience is supposed to find households on their way to spending more. Past upsold customers are now higher-tier buyers. A frequency growth audience is supposed to find occasional buyers before they become loyal ones. Past frequency converts are now loyal. In every case, the historical ground truth is a population whose relevant behavior has already changed, and what gets modeled is not what they looked like before, but who they are now.

So clients ask for proxies instead: active category buyers who aren’t yet brand buyers, households that recently entered the category, category switchers. These are buildable. They are also, in every meaningful sense, the wrong audiences. They approximate the thing the brand wanted without being it. Active category buyers include the brand’s future buyers but also people who will never buy it. Category switchers are interesting but not the same as pre-conversion signals. The proxy is not the audience that works. It’s the audience that exists.

2. What happens to the proxy next

A proxy audience is either targeted directly, if it’s large enough, or it becomes the seed for expansion, which is where the second gap opens.

The expansion step exists because a proxy audience is often too small to run a real campaign. Lookalike modeling takes the seed and extends it to a much larger population. What the extension actually finds depends entirely on what data is used for expansion, and here the industry has settled on an answer that makes the first problem worse instead of solving it.

Large third-party identity datasets used for audience expansion are built from static attributes, e.g., demographic data, psychographic profiles, purchase panel signals aggregated into categorical summaries. What they don’t contain is what a household was doing in the months before it made a specific behavioral change. What they contain is a snapshot of who a household is today, measured along dimensions that don’t actually distinguish pre-conversion households from anyone else.

When a third-party model expands a seed of active category non-buyers, it finds other households that share demographic and psychographic attributes with them. But the active category non-buyers in the seed are not a coherent population with distinctive attributes. They’re a proxy, defined by not having bought the brand yet, not by any behavioral pattern that predicts, with reasonable probability, that they will. Expanding them by attribute-matching produces more people who resemble the seed households by demographic profile, not more people showing the specific behavior that predicts conversion.

A more specific version of this failure is less obvious: even when the seed is correctly defined as past converters, households that actually became new-to-brand buyers in a historical window, the expansion step still fails in a particular way. All current brand customers became new-to-brand buyers at some point in the past. The majority of higher-tier buyers were upsold at some point. If you model a seed of past new-to-brand converters and ask who else looks like them, you find the demographics and psychographics of people who look like current brand buyers, because that’s what past new-to-brand buyers have in common: they became current brand buyers. The signal that was present before they converted is gone. The expansion doesn’t find people who are about to become your customers. It finds people who already resemble your current customers.

Strip away the machine learning and the result is familiar. It’s not meaningfully different from targeting women 25 to 35 with household income above $100,000 because that’s roughly what the current customer base looks like. The mechanism is more sophisticated. The targeting logic underneath it isn’t.

3. The semantic gap

The proxy problem and the expansion problem are both about the relationship between an audience definition and the households it actually finds. There is a third gap that runs alongside them, less about which households get selected and more about what signals are used to select them.

The signals available to a model trained on transaction data are limited to what the data contains and what the data’s structure allows the model to learn, e.g., category codes, merchant codes, purchase frequency, dollar amounts. These are useful, and a model trained on them does learn real patterns. However, they have a hard ceiling on what they can capture, and the ceiling matters for the specific task of finding pre-conversion households.

A model trained only on transaction data structured as category codes and merchant codes can only learn relationships that show up as co-occurrences in the data itself. If the relevant relationship is that gym-goers tend to buy sports nutrition, a statistical model learns this only if gym memberships and sports nutrition purchases co-occur frequently enough in the training data to be statistically notable. If the data doesn’t contain enough gym membership records, or if they’re coded in a way that obscures the relationship, the model never learns this relationship, not because it isn’t real, but because it never came up often enough to count.

A language model fine-tuned on real purchase outcomes starts from a different place entirely. It already knows, from the breadth of everything it was trained on before any fine-tuning, that a gym membership and a sports drink belong to the same story, and that buying a gym membership, running shoes, and a foam roller across different categories is the signature of someone building a fitness habit. “Lactose free” on a product description is a fact about the person buying it, not just the product, and private-label purchases signal price sensitivity in a way no category code encodes. Fine-tuning doesn’t teach the model about the world. It teaches the model to apply what it already knows to one specific, narrow business question, a much smaller and cheaper problem to solve. And critically, it does this automatically for every audience definition Geminus builds, without anyone having to identify the relevant cross-category connections in advance or encode them as rules. A sharp analyst could make the gym-to-sports-nutrition connection by hand for one brand in one campaign. The model makes it, and thousands of connections like it, across every brand and every campaign, automatically.

Running a frontier language model over tens of millions of households for every campaign would be prohibitively expensive. Geminus uses a teacher-student distillation approach developed with Understudy Labs: a frontier model establishes the standard on a sample, and a compact open-weight model is fine-tuned to reproduce the standard at a fraction of the cost. The compact model runs inside Geminus’s own infrastructure, on the full household universe, without sending household-level data to any external API.

The gap between statistical and language models is sharpest at the intersection of multiple data sources. A retail transaction panel records what was bought: a product, a brand, a category. A credit card transaction panel records where money went: a merchant, a type of business, across a much broader set of categories than any single retail panel covers. These two descriptions of household behavior live in completely different classification systems that were never designed to talk to each other. A household’s retail panel record might show heavy purchases in personal care while its card record shows heavy spend at premium grocery. A rule-based system treats these as two unrelated facts. A language model reading both as plain text describing the same household can synthesize them into a coherent picture: this is a household that invests in personal appearance and has the income to buy quality food. No one wrote a rule that encodes this connection, and no statistical model trained on a single panel can find a cross-panel signal like it. The full argument for why a language model closes this gap, and why a compact one makes more sense than a frontier API, is laid out in a companion post.

4. The measurement gap

Even if an audience is built correctly and the model is trained on appropriate signals, there’s the question of whether any of it actually caused the sales it’s credited with.

Attribution, the dominant measurement methodology in audience advertising, works by associating a post-exposure purchase with the campaign that preceded it. If a household was in the audience, saw the ad, and then bought the product, the sale gets attributed to the campaign. What attribution cannot determine is whether the household would have bought the product anyway, absent any targeting at all. The counterfactual is never observed.

This matters more than the industry acknowledges, because the households in a well-built audience are, by construction, the ones most likely to buy the product. That’s the point of the audience. But it also means that a measurable fraction of the sales attributed to the campaign would have happened regardless. Attribution assigns them to the campaign anyway. The headline lift number overstates the campaign’s actual contribution, and by how much is not observable from the data the measurement uses.

A randomized holdout doesn’t have this problem. Before the campaign activates, a random subset of the audience is set aside and never targeted. After the campaign, the targeted group’s purchase behavior and the holdout group’s purchase behavior are compared directly in the same transaction data used to build the audience. The lift is the observed difference between two groups that were identical except for one thing: whether they were shown anything. The difference is not what the campaign was present for. It’s what the campaign caused.

What an audience looks like when the gaps are closed

A Hill’s Pet Nutrition brand director wants to grow new buyers of the therapeutic pet food line, the kind a vet recommends for a specific condition, sold mostly through pet specialty retailers.

The objective goes to an AI agent, and the agent doesn’t ask what audience the brand wants. It reads the objective, works out what finding the right households actually requires, and identifies new-to-brand acquisition as the relevant strategy. What a positive outcome looks like gets defined precisely: a household that has purchased a competing therapeutic pet food brand but has not purchased Hill’s in the same period. The agent queries that directly from the transaction panel, finding households that match across all available retail coverage, not just the pet specialty channel but grocery and mass retail too, because that’s where some therapeutic pet food buyers happen to shop. It runs automatically, and the brand director sees the formal definition before any scoring starts and can challenge it.

The model is trained on the real historical record of this strategy, what households looked like before they became new-to-brand Hill’s buyers, not after. It scores the entire eligible universe. PetSmart can find its own non-buyers. Chewy can find its own non-buyers. Neither can see across into the other’s transactions, and neither can find the household buying a competing therapeutic brand at a grocery store across town. Geminus can, because the panel is cross-retailer.

A holdout group is set aside before activation. The treatment group is targeted through the brand’s activation platform of choice. Weeks later, the comparison is direct: how many more households in the targeted group started buying Hill’s than in the holdout group, measured in the same transaction data used to build the audience. The number is what the campaign caused.

None of this is a set of separate features. It’s one continuous design choice: change the objective and the definition changes, change the definition and the score changes, change the score and the holdout result means something different. Pull one piece out and the rest stops meaning anything. The audience market sells proxies because the alternatives were too expensive, too slow, or structurally impossible. They’re not anymore. That’s the whole pitch. What each of these design choices means in practice is the subject of a separate piece on what Geminus audiences actually are.

vadim@geminusdata.ai