Why Your Fraud Detection AI Is Measuring the Wrong Baseline
[custom_breadcrumb]
Home > Blog > Your Fraud AI Is Measuring the Wrong Baseline That Is Why Sophisticated Financial Crime Is Getting Through

There is a question most banking technology leaders have not been asked directly about their fraud detection AI. Not ‘is the model performing?’ – that one gets asked constantly, with accuracy metrics produced in response. The question that matters more is different: what is the model measuring performance against?

The answer to that question reveals more about a fraud detection programme’s actual effectiveness than any accuracy benchmark. Because the baseline the model uses – the reference against which it decides whether a transaction is anomalous – determines not just how many fraudulent transactions it catches, but which ones it misses. And the ones it misses are not random. They are precisely the sophisticated schemes that were designed to avoid detection.

The Baseline Problem – Rules-Based vs Per-Entity

Rules-based fraud detection was built on a structural assumption: fraudulent transactions share identifiable characteristics that can be encoded as thresholds. Amount above X. Velocity exceeding Y. Geographic distance from Z. The assumption was correct – for the fraud typologies that existed when the rules were written. 

The problem is not the assumption. It is that the assumption is static and fraud is adaptive. Bad actors observe detection thresholds and engineer their activity to stay below them. A transaction that is individually unremarkable by every categorical measure can be part of a sophisticated scheme operating across months and multiple jurisdictions – invisible to a system that asks only ‘does this transaction match a known fraud pattern?’ 

AI changes the question. Instead of asking whether a transaction matches a predefined typology, AI asks: does this transaction deviate from the behavioural baseline of the specific account it originates from? 

This is the difference between a segment baseline and a per-entity baseline. A segment baseline sets a threshold that applies to every account in a category. A per-entity baseline reflects the specific behavioural history of this account, this customer, this relationship with the bank. Two transactions of identical value, identical merchant category, and identical timing can produce opposite outcomes from a per-entity AI system – one flagged, one not – because one deviates from its account’s individual history and the other is entirely consistent with it.

The question is not ‘does this transaction look like fraud in general?’ It is ‘does this transaction look like fraud for this customer, given everything we know about how this account has behaved over time?’

The Scale of the False Positive Problem

The-Scale-of-the-False-Positive-Problem-Maveric-Systems

These figures are not primarily a customer experience problem, though false positives at scale certainly create one. They are an efficiency problem that undermines the entire efficiency case for fraud automation. A fraud system generating false positives at 95–98% rate is not reducing operational cost. It is replacing the cost of human review with the cost of AI-generated alert investigation at volume – and in many cases at higher total cost, because the investigation infrastructure required to process AI-generated alerts is more expensive than the human review it replaced.

Per-entity baseline monitoring reduces false positives dramatically – because the baseline against which each transaction is assessed is far more sensitive to genuine anomalies and far more specific to normal behaviour for that entity. The same model that generates 97 false alerts per genuine suspicious transaction on a segment baseline generates substantially fewer on a per-entity baseline, because the definition of ‘anomalous’ is calibrated to each individual account rather than to a category average.

What Per-Entity Monitoring Catches That Segment Monitoring Cannot

The fraud typologies most resistant to rules-based detection share a structural characteristic: they are designed to look unremarkable in any single dimension while being deeply anomalous in the context of the specific account they target. 

A corporate account receiving seventeen international wire transfers in a month is unremarkable if the segment average is high. It is highly anomalous if this specific corporate account has received two per month consistently for three years. The segment rule does not see this. The per-entity baseline does. 

Temporal patterns that span years. Geographic distributions that look normal in isolation but anomalous over a multi-year relationship. Incremental changes in counterparty behaviour that individually trigger nothing but collectively represent a systematic shift in account activity. All of these are visible at the per-entity level and invisible at the segment level – which is precisely why sophisticated schemes are designed to exploit the segment-level blind spot. 

AI enables the assessment of these patterns because it can maintain and reason over a behavioural history that is specific to each entity and extends across the full relationship with the bank. Rules-based systems were not limited by concept – the statistical techniques existed. They were limited by computing power. The per-entity, multi-year, multi-geography assessment that AI makes possible was not computationally feasible at production scale before AI. It is now.

THE DRIFT RISK AT PRE-ENTRY SCALE
Per-entity baseline monitoring introduces a governance requirement that programme-level monitoring does not: the baseline itself must be maintained. A customer whose legitimate behaviour changes – a business that expands internationally, a consumer whose income and spending patterns evolve – will generate false positives if the baseline does not update to reflect the change. Governance of per-entity baseline models requires monitoring at the same granularity as the model operates, and data contract governance for every upstream feed the baseline depends on. The per-entity precision that makes the model accurate depends entirely on the data consistency that keeps the baseline current.

The Question Your Programme Should Be Answering

If 95-98% of your fraud alerts are false positives, what is that telling you about the baseline your model is working from?

It is telling you that the model is measuring transactions against a standard that is too generic to distinguish anomalous from merely unusual. That the segment threshold you are using is set at a level that most unusual transactions exceed but only a small fraction of fraudulent ones uniquely trigger. That the investigation overhead you are managing is the operational cost of a baseline problem, not a model problem.

The shift from segment-level to per-entity baseline monitoring is not a model upgrade. It is an architectural change in what the model is measuring. And it is the change that moves fraud detection from a volume problem – how many alerts can we process? – to an intelligence problem – how accurately can we distinguish genuine anomalies from normal behaviour in this specific account?

The institutions that have made this shift do not have more accurate models. They have more precisely calibrated baselines. The model is the same technology. What it is measuring is fundamentally different.

The full architecture – per-entity behavioural baseline monitoring, continuous drift detection at entity-segment level, data contract governance for upstream feeds, and the governance framework for AML alert triage – is set out in CIO Mandate Series Paper 3.
Download: Engineering Trusted Automation in AI-First Banking Operations

Article by

Maveric Systems