Organizing Brownfield Data Across Multiple Plants.
The Missing Relationship Problem in Financial Services AI
The risk model flagged the exposure correctly. It just didn't know the exposure it flagged and the exposure sitting in a different business unit belonged to the same corporate family. (This is an illustrative scenario reflecting a pattern we see across financial institutions, not a specific customer engagement.)
A risk that was visible in two places and connected in neither
Consider a commercial bank whose credit risk model correctly flags a deteriorating exposure to a mid-market manufacturer. Around the same time, the bank's trade finance desk is extending a separate facility to what its own systems record as a different counterparty. Both positions are accurate, well-governed, and individually within policy. What neither system knows is that the manufacturer and the trade finance counterparty are subsidiaries of the same parent company — a relationship that exists in a legal entity database maintained by a different team, for a different purpose.
This kind of gap is common in financial services, and it rarely traces back to a data quality problem. The records are accurate. The models are sound. What's missing is the relationship connecting two accurate pictures into one more complete picture of exposure — a relationship that lives in a system neither the credit model nor the trade finance desk had any reason to query.
Three places the gap shows up
This pattern tends to recur in a few recognizable forms across banks and financial institutions running risk, compliance, and client analytics on Databricks.
1. Collateral pledged more than once across a corporate group
The operational question: "This asset secures this facility. Is it also pledged against another facility elsewhere in the corporate group, and what happens to our overall secured position if its valuation changes?"
Answering that means tracing a chain that spans several entities: the borrower, the facility it secures, the collateral itself, any guarantee attached to it, and any related borrower within the same corporate group who may have pledged the same or related collateral against a separate facility. Each individual facility and its collateral can be recorded accurately and still leave the bank's overall secured position unclear, because the relationship connecting one facility's collateral to another borrower's facility elsewhere in the group typically isn't visible to either loan officer managing their own piece of it.
2. Aggregate exposure across a corporate family tree
The operational question: "What is our total exposure to this corporate family, once we account for every subsidiary, every guarantee, and every facility extended by each of our business units separately?"
Answering that well requires connecting a parent entity to every subsidiary it controls, every subsidiary to whatever facilities and guarantees it holds across the bank's different business units, and every one of those facilities to its current exposure. That legal entity hierarchy typically originates in systems built for KYC, regulatory, or onboarding purposes, separate from the systems where credit and trade finance decisions actually get made, which is exactly how two business units can each stay within policy individually while a more complete view would show meaningfully greater exposure to one corporate family than any single view reveals.
3. The beneficial ownership chain that AML monitoring can't fully see
The operational question: "This account's activity looks unusual. Before we escalate, can we confirm whether the ultimate beneficial owner is connected to any other account we've already flagged?"
AML and transaction monitoring systems are generally very good at detecting unusual activity within a single account. Tracing that account back through layers of beneficial ownership to a holding structure, and from there to any other account connected to the same ultimate owner, often depends on a separate ownership registry that the monitoring system wasn't built to query directly. A genuinely connected pattern across two accounts can go unnoticed simply because the relationship between them wasn't available to the system doing the monitoring.
Why none of this reflects a shortfall in the underlying data or models
In every one of these situations, the underlying records are accurate, and the systems holding them are generally well-governed for their own purpose. But that data rarely lives in one place. Customer risk data may sit in a core banking platform. Legal entity hierarchies often originate in KYC or regulatory onboarding systems. Beneficial ownership registries may come from a third-party data provider entirely. That the relevant entities and relationships are spread across core banking, KYC, AML, MDM, and regulatory systems isn't unusual, it's the normal state of a financial institution's technology estate.
The data was governed in every system that held it. The relationship connecting those systems existed in none of them.
Unity Catalog and the Databricks platform provide the governed foundation for bringing that landscape together. The next challenge for most financial institutions running risk and compliance workloads at scale isn't more governance over any one source system — it's making the relationships between entities across that landscape reusable, so a credit model, a trade finance system, and an AML monitoring platform can all draw on the same connected understanding of corporate structure and ownership.
Connecting corporate structure, collateral, and ownership
Kobai helps financial institutions connect these relationships once, as Shared Business Context that brings together legal entity hierarchies, collateral and guarantee relationships, and beneficial ownership chains from wherever they originate (such as core banking, KYC, AML, or third-party registries) into a single connected model on Databricks. That context becomes available to the models, applications, and workflows that perform credit, trade finance, and AML functions, rather than replacing them: a credit model can reference the same corporate family relationships, a trade finance system can check the same collateral chain, and an AML workflow can draw on the same beneficial ownership structure.
Making this context available consistently changes what these workflows can see. A more complete view of aggregate exposure to a corporate family becomes available, rather than requiring manual reconstruction across business units under deadline pressure. AML investigators can check related accounts as a matter of course, rather than depending on someone remembering to query a separate registry. And credit decisions can be made with visibility into collateral relationships that span more than one facility, rather than each loan officer seeing only the piece they manage directly.
|
Without Shared Business Context |
With Databricks + Kobai |
|
Collateral pledged across multiple facilities isn't visible to any single loan officer. |
Collateral and guarantee relationships across facilities are available to credit workflows. |
|
Aggregate exposure to a corporate family is reconstructed manually across business units. |
A more complete view of aggregate exposure is available from connected entity relationships. |
|
AML escalations depend on an investigator manually checking a separate ownership registry. |
Beneficial ownership relationships are available to support the escalation review. |
|
Two business units can each stay within policy while overall exposure is greater than either view shows. |
Overall exposure is more visible earlier, before it becomes a policy or regulatory concern. |
Where years of data investment leaves a gap
Most financial institutions have invested seriously in data infrastructure over the past decade — consolidating records, governing access, and building genuinely capable risk and compliance models on top of that foundation. That investment shows up clearly in how well each individual system performs at the job it was built for.
One reason that investment hasn't always translated into AI that scales confidently across risk, compliance, and client management is that the relationships connecting entities across those systems were never anyone's job to build. Each function's data was accurate. The picture that mattered — how those accurate pictures fit together — never fully existed anywhere.
Aggregate risk isn't always visible in any single dataset. It can emerge from the relationships between entities, exposures, and ownership structures that individual systems only see in part.

