Skip to content
1*zn2FQFJ5Fq_MLIu9-zISzA-2-200x200
Semantic Distillation: A Brief Primer

The fact that business teams are drowning in disconnected data is getting to be a bit of a cliche. Adding a semantic layer to an enterprise data platform can bring order to chaos, allowing teams to collaborate effectively and leverage AI to unlock valuable insights.

Celebal Technologies Partners with Kobai
Celebal Technologies Partners with Kobai

to Launch Turnkey Knowledge Graph Solutions For Global
Enterprises on Databricks

Latest Event:
Webminar on Wednesday, October 29th, 2025
Play now
Why Cross-Sell AI Fails in CPG Without Shared Business Context
KobaiSep 2, 2026, 2:51:08 AM6 min read

Why Cross-Sell AI Fails in CPG Without Shared Business Context

Why Cross-Sell AI Fails in CPG Without Shared Business Context
7:41

The recommendation engine suggested the bundle. The retailer couldn't actually sell it. Nobody caught the mismatch until the promotion was already live.

A recommendation that looked right and wasn't

Picture a cross-sell model at a large CPG manufacturer recommending a snack line paired with a top-selling beverage at a major retailer's stores in one region. The affinity itself isn't in question — customers who buy one genuinely tend to buy the other. But the beverage has just been reformulated into a new package size that hasn't rolled out to that retailer's distribution centers in the region yet. The promotion launches, in-store displays go up, and a meaningful share of the recommended bundle simply isn't on the shelf where the promotion is running. (This is an illustrative scenario built to reflect a pattern we see across CPG cross-sell programs, not a specific customer engagement.)

This kind of mismatch is common in CPG cross-sell and bundling programs, and it rarely shows up as a data quality problem. The purchase history is accurate. The affinity model is sound. What breaks down is everything sitting between "these products sell well together" and "this specific customer, at this specific retailer, in this specific region, can actually buy both right now" — a chain of business relationships that the recommendation model was never built to check.

 

Three places the gap shows up

This pattern tends to recur in a few recognizable forms across CPG organizations running cross-sell and trade promotion programs on Databricks.

1. The SKU transition that the business understands and the AI doesn't

The operational question: "We just replaced this SKU with a new package format. The old SKU has three years of purchase history the recommendation model relies on. The new SKU is what's actually in distribution. Does the model know they're the same product?"

A SKU transition like this is a completely ordinary part of running a CPG business, and everyone involved understands the relationship intuitively: the old identifier and the new one refer to the same underlying product, just at different points in its lifecycle. But that relationship typically isn't explicit anywhere a recommendation model can see it. The retailer's assortment system may still reference the old SKU. The promotional calendar may reference the new one. The recommendation model, trained on historical purchase data, may be reasoning entirely from an identifier that no longer has any inventory behind it, recommending a product that, as far as distribution is concerned, doesn't exist.

2. The bundle that's only valid in some regions

The operational question: "Can we recommend this bundle to shoppers at this retailer, given which distribution center serves their region, what's currently in that center's inventory, and which promotion is active there this week?"

Answering that well means connecting several things that typically live in different systems: which retailer and region a shopper belongs to, which distribution center serves that region, what that center currently has in stock, and which promotional calendar applies there this week. A recommendation model trained on national purchase patterns has no inherent way of knowing that a bundle valid in one region is meaningless in another this month, because the underlying distribution and promotional relationships were never connected to the recommendation itself.

3. The bundle that violates a trade agreement nobody checked

The operational question: "Does recommending this bundle put us in conflict with a category-exclusivity clause we signed with this retailer for a competing brand?"

Trade agreements and retailer-specific exclusivity clauses are usually negotiated and tracked by the commercial team responsible for that relationship, typically in a contract management or TPM system built for that purpose, not in the systems that generate cross-sell recommendations. A model with no visibility into that relationship can recommend a technically appealing bundle that quietly conflicts with a contractual commitment, and the first anyone hears about it is when the retailer's category manager raises it.

 

Why the recommendation engine isn't the problem

In each of these situations, the underlying data exists and is generally well-managed for its own purpose. But it rarely lives in one place. Purchase history and recommendation models typically run on Databricks; distribution and inventory data may sit in SAP; trade agreements are often tracked in a separate contract management or TPM system entirely. That the relevant context is spread across several systems isn't an edge case. It's the normal state of a CPG technology estate, and it's exactly why no single system, including the recommendation engine, was ever positioned to see the full picture on its own.

Unity Catalog provides the governed foundation for whatever data does live on the Databricks Lakehouse. The next challenge for most CPG organizations running cross-sell and trade programs isn't more governance over any one system. It's making the business relationships that span all of these systems reusable, so a recommendation engine, a promotional planning tool, and a trade compliance process can all reason from the same connected understanding of product identity, distribution, and commercial agreements.

 

Connecting product identity, distribution, and trade context

Kobai helps CPG organizations connect these relationships once, as Shared Business Context that draws together product identity across SKU transitions, retailer and regional distribution relationships, and trade agreement terms from wherever they live (Databricks, SAP, or a dedicated contract system) into a single connected model. That context becomes available to whatever system needs it: a recommendation engine can reference the same product-identity relationship that resolves an old SKU to its replacement, a promotional planning tool can draw on the same distribution and inventory relationships, and a compliance review can check the same trade agreement terms.

The commercial value of making this context available consistently shows up directly in how these programs run. A recommendation engine that can reference resolved product identity is far less likely to recommend a SKU that no longer has inventory behind it. A promotional planning process with visibility into regional distribution can catch a mismatch before a campaign launches rather than after. And a trade compliance review that has commercial agreement terms available alongside the proposed bundle can flag a potential conflict during planning, rather than after a retailer's category manager raises it.

 

Without Shared Business Context

With Databricks + Kobai

A SKU transition leaves the recommendation model reasoning from an identifier with no inventory.

Product identity across a SKU transition is resolved and available to the recommendation engine.

A national model has no visibility into regional distribution or inventory relationships.

Regional distribution and inventory relationships are available to the recommendation workflow.

Trade exclusivity terms live in a separate system from cross-sell and promotion tools.

Trade agreement terms are available alongside the proposed bundle during planning and review.

Underperforming promotions are diagnosed after the fact, often too late to correct.

Relevant context is available earlier in the process, before a campaign launches.

 

From identifying an opportunity to acting on it

Cross-sell gets much more powerful when AI can reason beyond what customers are likely to buy together and understand what they can actually buy — at this retailer, in this region, this week, under the commercial agreements already in place.

That's the difference between identifying an opportunity and being able to act on it.

COMMENTS

RELATED ARTICLES