There is a specific moment, recognizable across a huge share of banking transformation programs, when a well-designed initiative runs headfirst into a wall that has nothing to do with its architecture, its strategy, or its funding: the data it depends on turns out to be fragmented, inconsistent, or simply not trustworthy enough to support the decision the new capability is supposed to make. This is data debt, and it is, at its root, a data management in banking problem inherited from years of prior systems and decisions, not created by the current program at all, which is exactly why it rarely appears on a program’s original risk register.
The Cost Is Larger Than Most Programs Budget For
Gartner’s research puts the average annual cost of poor data quality to an organization at $12.9 million, a figure that spans lost revenue, wasted operational effort, and flawed decisions built on unreliable inputs. That figure is an average across all industries; in banking specifically, the cost tends to run higher, because weak data management in banking does not just create inefficiency it creates regulatory exposure, through inaccurate reporting, failed reconciliations, and AML or KYC gaps that a poorly governed data estate makes far more likely.
What makes this figure particularly relevant to transformation programs is that it is a recurring, ongoing cost not a one-time cleanup expense. A bank carrying significant data debt pays this cost every year the underlying data problem goes unaddressed, whether or not a transformation program is actively running. A transformation program that does not address data management in banking as part of its scope is not avoiding that cost. It is simply choosing to keep paying it, on top of whatever the new initiative itself costs to build.
Why Data Debt Gets Discovered Mid-Program, Not Before It
Data debt is structurally difficult to catch during initial planning, for a specific reason: most data quality assessments happen at a sample or summary level, and fragmentation problems tend to be local concentrated in specific product lines, specific legacy systems, or specific customer segments that a high-level assessment does not surface. A program can pass an initial data readiness check and still discover, once it starts building against the full production data set, that a meaningful subset of records are duplicated, inconsistent, or simply wrong in ways the original assessment did not catch.
This is precisely why data debt shows up so consistently as a mid-program discovery rather than a pre-program red flag. By the time it surfaces, a program has usually already committed significant build effort to an architecture that assumed cleaner data than what actually exists, meaning the fix is not just cleaning the data, but reworking the parts of the build that assumed it was already clean.

Where This Hits Hardest: Banking Customer Data Management, AML, KYC, and Regulatory Reporting
Nowhere is data debt more consequential, or more visible to a regulator, than in banking customer data management. Most institutions still hold a customer’s information across a patchwork of siloed systems core banking, cards, lending, wealth, digital channels each with its own version of the same person, updated on its own schedule, rarely reconciled against the others. Without a single, governed, unified customer data record, every downstream initiative that depends on knowing who a customer actually is, consistently and in real time, inherits that fragmentation.
Anti-money laundering detection is one of the clearest illustrations of how this plays out. AI-driven transaction monitoring and alert investigation tools are only as effective as the data feeding them and duplicate customer records, inconsistent transaction categorization across channels, and incomplete data captured at onboarding all directly degrade an AML system’s ability to detect genuine risk, regardless of how sophisticated the detection model itself is. The same pattern repeats in regulatory reporting, where reconciliation exceptions and inconsistent reference data across systems routinely turn what should be a straightforward reporting exercise into a manual, error-prone reconciliation effort that consumes far more resources than the reporting requirement itself would suggest.
In both cases, the model or system being built is frequently blamed for underperformance that actually originates in the data layer beneath it leading institutions to invest in more sophisticated detection technology when the more effective fix would have been addressing the customer data management gap the technology was never going to be able to compensate for.
What This Looks Like When Banks Actually Fix It
The remediation pattern shows up consistently wherever institutions have taken data management in banking seriously as a first-class workstream rather than a background cleanup task. In one global tier-one bank’s legacy modernization effort, years of undocumented data mapping and transformation rules had accumulated across multiple legacy application stacks nobody could say with confidence what a given transformation rule actually did, or why, which made data quality validation largely guesswork. Rather than attempting to rebuild the rules from scratch, the bank used generative AI to reverse-engineer and document the existing rules directly from the legacy code and data flows themselves, turning years of undocumented institutional assumption into an auditable, governed rule set a modernization program could actually build on.
A separate initiative at another tier-one bank took a similar approach to the audit side of data governance: moving from static, sample-based data quality audits – manual, infrequent, and inherently unable to catch most issues between audit cycles, to an AI-enabled, adaptive data quality ecosystem capable of continuous, business-aligned validation across business, logical, and technical rules simultaneously. Both examples illustrate the same underlying principle: data debt, once made visible and continuously monitored rather than periodically sampled, stops being an open-ended liability and becomes a governed, remediable asset a transformation program can actually depend on.
The Case for a Single, Governed View of the Customer
A recurring theme across both examples above, and across most successful data management in banking initiatives more broadly, is consolidation. Institutions that invest in master data management and a genuine single customer view one governed record that every system references, rather than each system maintaining its own consistently spend less time firefighting downstream data quality incidents than institutions that keep patching the symptoms of fragmented customer data one integration at a time.
This matters more, not less, as AI takes on a larger share of customer-facing and risk decisions. A model recommending a product, flagging a transaction, or approving a credit line is only as reliable as the unified customer data behind it. Banking customer data management that remains siloed by product line or channel does not just create operational friction it directly caps how much a bank can trust any AI system built on top of it, because the AI is only ever as good as the least reliable data source it draws from.
Why Treating Data Readiness as a Parallel Workstream Does not Work
The most common mistake in program design is scheduling data quality remediation as a workstream that runs alongside the main build, on the assumption that both can converge by go-live. In practice, this sequencing almost guarantees the wall gets hit mid-program rather than avoided: the build workstream makes architectural assumptions about data availability and quality before the data workstream has actually delivered on those assumptions, and when the two do not align which, given the scale of typical data fragmentation, they usually do not the build has to be reworked around whatever data reality the remediation effort actually produces.
The more resilient approach, visible in the programs that avoid this wall, treats data readiness as a genuine prerequisite rather than a parallel track: assessing data quality for the specific systems and use cases in scope before architecture decisions are finalized, and sequencing the build to depend only on data that has already been validated as ready, rather than data that is assumed to be ready by the time it is needed.
What “AI-Ready” Data Actually Requires
For programs specifically involving AI which is now most of them the bar is higher than general data quality. AI-ready data needs to be not just accurate, but governed, lineage-tracked, and contextualized: an AI system making a consequential decision needs to be able to show where its input data came from and how it was validated, not just that the data happened to be correct. This is the layer beneath every trust principle a bank applies to its AI systems a model cannot be explainable if the data behind its decision cannot itself be explained.
Institutions that treat data management in banking and AI investment as two line items on the same budget rather than the data investment being a prerequisite that unlocks the AI investment are consistently the ones that avoid rediscovering the same wall on their next initiative. The $12.9 million average annual cost of poor data quality does not go away because a bank launched an AI program on top of it. It compounds.
Related reading:
- Why Most Banking Transformations Fail at Execution (pillar guide)
- How to Improve Data Quality in Banking Without Breaking Systems
- AML Data Challenges: Why Detection Fails Despite Investment
Sources
Gartner research on the average annual cost of poor data quality ($12.9 million)
FAQ
1. Why does data debt usually surface mid-program rather than during initial planning?
Most data quality assessments happen at a sample or summary level, while fragmentation problems tend to be local, concentrated in specific product lines or legacy systems that a high-level assessment does not surface until the program starts building against the full production data set.
2. What does ‘AI-ready’ data require beyond general accuracy?
Data has to be governed, lineage-tracked, and contextualized, meaning an AI system making a consequential decision needs to show where its input data came from and how it was validated, not just that the data happened to be correct.
3. Why does treating data readiness as a parallel workstream usually fail?
The build workstream makes architectural assumptions about data availability before the data workstream has delivered on those assumptions. Given the scale of typical fragmentation, the two rarely align by go-live, and the build ends up reworked around whatever data reality the remediation effort actually produces.