Almost everything sold today under the name “audience” is a guess dressed up as data. A pre-built segment is a shelf item: a fixed definition of frequent shoppers or high-income households, chosen not because it matches what the brand is trying to accomplish but because it exists. A custom model is a project: weeks of analyst time translating a business objective into a technical spec, then building against it. Neither starts from what the brand actually wants to do. Neither proves, at the end, that it worked.

THE TARGETING GAPTargeted todayBroad proxy audienceHouseholdsthat matteron a real trajectorytoward purchaseBudget wastedon wrong householdsRight householdsmissed entirely

We start from the objective itself. A brand states a growth goal in plain language, e.g., grow the category, launch a new product, grow how much customers spend with the brand, or retain customers and prevent churn. A Geminus AI agent reads the goal and does the translation work, which used to require a strategist and an analyst: it produces the formal definition of what a positive outcome looks like, who is eligible, and over what window. This happens before any model is trained, and the definition is available for a person to review and challenge before any household is scored.

Everything we build from the definition runs on real, observed purchases: what a household bought, from which merchant, for how much, when, tracked across retailers in a panel covering tens of millions of households. The underlying data is not an ad exposure used as a proxy for interest, or a survey extrapolated to a population, or a modeled third-party segment several hops removed from anything anyone actually did. Every layer of the pipeline, the model’s accuracy, the scores households receive, the lift measured afterward, is only as trustworthy as what sits underneath it. Here, it actually happened.

The model that scores the households is a compact language model we fine-tune on real purchase outcomes, and it finds things statistical models trained on a single company’s own data cannot reach. A statistical model can only learn relationships that show up often enough in its training data to count, and if the connection between two behaviors isn’t common enough in one company’s records, the model will never find it regardless of whether the relationship is real. We start instead from a model that has read the world: it already understands that a household buying a gym membership, running shoes, and a foam roller is someone building a fitness habit who hasn’t yet crossed into sports nutrition, without needing to count the combination in the panel. The model already knows it. We fine-tune the existing understanding into a calibrated answer to one specific business question, and it applies automatically to every audience we build, without anyone having to identify the relevant connections in advance or encode them as rules. Running a frontier language model over tens of millions of households for every campaign would be prohibitively expensive. So we use a teacher-student distillation approach instead: a frontier model establishes the standard on a sample, and the compact model we deploy is fine-tuned to reproduce it at a fraction of the cost, running inside our own infrastructure on the full household universe. CPG audience-building is the first vertical AI model Geminus is building. The same architecture applies to any domain where the gap between what a business needs and what current tools can produce is large enough to be worth solving.

We build proof into the design from the start. Before a campaign activates, Geminus automatically sets aside a random subset of the scored audience, a holdout group that’s never targeted. After the campaign, the holdout group’s and the targeted group’s purchase rates are compared directly in the same transaction data used to build the audience. The lift is the observed difference between the purchase rate in two groups that were identical in every respect except one: whether they were reached by the campaign at all.

Status quo Geminus
Starting point Client specifies the audience Client states the business objective
Audience definition Fixed segment, or a spec the client writes AI agent translates objective to formal definition
How it’s built Scoped project, weeks, one analyst Automated, days, reusable
Underlying data Often inferred: exposure logs, surveys, modeled third-party segments Real, observed transactions
Modeling Rules, or statistics on a company’s own data World knowledge plus machine learning, fine-tuned on real outcomes
Proof Attribution, a modeled estimate Randomized holdout, observed lift

The full argument for why proxy audiences and lookalike expansion fall short is in The audience you’re buying is not the audience you need.

vadim@geminusdata.ai