The governance conversation around AI compliance in banking tends to operate at the wrong level of precision. It asks whether institutions have model validation frameworks, whether governance documentation exists, whether oversight processes were followed. These are the right questions for the previous era. For the AI-first compliance function, they are necessary but not sufficient.
The questions that reveal whether an institution’s compliance AI is genuinely governed – rather than documented – are more specific. They go to the granularity at which the AI actually operates, and whether the governance architecture has kept pace with that granularity.
Seven of those questions are set out below, with the answers that the depth of AI capability in banking now enables. They are the questions a technically informed board member, a demanding examiner, or a CTO briefing a CIO should be asking. Most institutions cannot answer all seven with confidence. The gap between confidence and uncertainty is the governance gap that the next examination cycle will find.
THE STANDARD THAT HAS MOVED:

Regulators are no longer asking ‘Do you have governance documentation?’ They are asking: ‘Can you demonstrate, using production data, that your compliance AI is performing within its defined parameters today?’ These seven questions are the operational expression of that standard.
QUESTION #1: In your AML system, does determination-level explainability include the semantic content of transaction narratives – or only structured fields?
The answer that matters: in an AI-first AML system, determination-level explainability is richer than feature attribution in the traditional sense. The model can assess the meaning of what is written in the payment description, the counterparty reference, the memo field – not just the structured fields surrounding it.
A rules-based system could detect the presence of a flagged word in a narrative. It could not assess whether the narrative as a whole made sense in the context of the transaction. AI can. The determination can therefore include not just “these structured features contributed” but “the language pattern in this transaction description was semantically inconsistent with the stated purpose of the payment.”
This is the explanation a skilled compliance analyst makes when reading a transaction manually. AI makes it at scale and consistently. An AML system designed to surface this explanation as a first-class output alongside every alert gives the compliance analyst something to exercise professional judgment on. An AML system that produces a confidence score gives them something to interpret blindly.
If your answer to this question is “our system traces to structured features,” your SAR narrative process is carrying a risk your examiners have begun to look for.
QUESTION #2:Does your AML monitoring establish a behavioural baseline at the individual entity level – or at the segment level?
This is the distinction between AI-first monitoring and rules-based monitoring stated precisely. In a rules-based system, the threshold was a category rule – a transaction amount, a frequency, a counterparty count – applied uniformly to all entities in a segment. The baseline was not a baseline for any specific entity. It was a threshold that applied to everyone in the group.
AI establishes a behavioural baseline for each individual entity it monitors. This account’s typical transaction frequency. This corporate’s expected counterparty geography. This customer’s normal payment patterns. Deviation from that entity’s baseline is what triggers the determination – not deviation from a segment average.
The consequence is significant in both directions. Sophisticated financial crime that operates just below aggregate thresholds is visible at the entity level – it deviates from this account’s own pattern even when it would not trigger any segment rule. And legitimate changes in customer behaviour must be reflected in baseline updates, or the model generates false positives for normal activity that has simply evolved.
Per-entity baseline monitoring is a more sensitive detection capability and a more demanding governance discipline simultaneously. If your answer is “we monitor at the model or segment level,” the adversarial financial crime that adapts to aggregate thresholds is the exposure your current architecture cannot see.
QUESTION #3:Can your credit adverse action notice trace the validation logic to the specific passage in the credit policy that governed the decision?
Most institutions can trace an adverse credit decision to the data inputs and the model weights. Fewer can trace it to the governing authority – the specific passage in the credit policy that the decision reflects.
The distinction matters for a reason that goes beyond regulatory compliance. When a regulator or a customer challenges an adverse credit decision, “the model considered these factors and weighted them as follows” is an explanation of mechanics. “This decision reflects the credit policy’s requirement that applications with revolving utilisation above 90% are assessed at a higher risk tier – and this application met that condition” is an explanation of authority.
AI makes the second explanation possible at scale through sentence-level lineage – tracing the validation logic applied to a decision back to the specific passage in the governing document that informed it. This makes AI-first credit compliance more defensible than human underwriting, not less. A human underwriter’s decision traces to their judgment. An AI decision with sentence-level policy lineage traces to the policy itself, which is documented, versioned, and auditable.
If your adverse action infrastructure traces to data inputs but not to policy passages, the explainability architecture is half-built.
QUESTION #4: Does your fair lending bias assessment cover unstructured data inputs?
Standard disparate impact testing is calibrated for structured features – income, credit score, utilisation rate, employment status. It asks whether the model’s outcomes differ across protected demographic groups controlling for these structured credit risk factors.
In an AI-first credit function that reasons over unstructured data – application narratives, customer correspondence, document text – this standard is necessary but insufficient. Language patterns in application text can correlate with demographic characteristics in ways that structured fields do not capture. The way applicants describe their employment situation, their housing circumstances, their purpose for the loan – these carry signals that can proxy for protected class membership even when demographic fields are explicitly excluded.
An AI model trained on data that includes unstructured text learns these correlations unless the training process explicitly identifies and addresses them. Standard disparate impact testing will not detect this exposure because the relevant features are not named. They are patterns in language the model learned implicitly.
Fair lending governance in an AI-first credit function must extend to the unstructured training data surface. If your bias assessment covers structured features only, the AI’s exposure to encoded demographic signals in text is ungoverned.
QUESTION #5: When your AI generates a regulatory submission, does the lineage for each figure trace to the specific passage in the governing regulatory standard?
Regulatory report data lineage typically traces figures to data sources and transformation steps. AI-first regulatory reporting makes a further level of tracing possible: from the figure, through the validation logic applied to it, to the specific paragraph in the regulatory standard that the validation reflects.
This is not a theoretical capability. The architectural condition that makes it real is deliberate design in the query-level validation layer. When an AI system generates a validation rule for a regulatory report figure, the prompt that generates it can include a citation to the specific regulatory requirement it reflects the precise principle, the paragraph number, the specific obligation being met. That citation becomes part of the validation rule’s metadata. The lineage is not traced retrospectively. It is built into the generation process.
The practical consequence: a regulator examining a submission can receive, for any figure they question, not just “here is where the data came from” but “here is the specific BCBS 239 principle and paragraph that this figure is designed to satisfy, and here is the validation logic that tested whether it did.” This is not just more efficient regulatory reporting. It is a categorically different level of regulatory defensibility – and it requires building the query-level validation architecture with regulatory citation metadata from the point of construction.
If your regulatory reporting lineage traces to data sources but not to regulatory passages, the submission is automated but not fully auditable.
QUESTION #6: Does your compliance resilience framework’s predict function monitor at the context-specific level?
Predict-prevent-mitigate-recover-restore is the right operational resilience framework for AI-first compliance. The question is at what level the predict function operates.
Programme-level predict monitoring asks: is this compliance AI programme performing within its governance parameters? Context-specific predict monitoring asks: is it performing within its governance parameters for this customer type, this transaction type, this market condition, this regulatory context?
These are different questions and they surface different failures. A programme-level metric can show acceptable aggregate performance while a specific customer segment is being systematically underserved or overscrutinised – a failure that is invisible in the aggregate and visible only at the context-specific level. Sophisticated financial crime adapts to programme-level monitoring and stays below aggregate thresholds. It does not adapt to per-entity context-specific monitoring because each entity’s baseline is individual.
The predict function in an AI-first compliance estate monitors not just “is the AML model performing?” but “is the AML model performing for wire transfers from newly onboarded corporate accounts in high-risk jurisdictions?” – the context where failure is most consequential and most likely to be exploited. If your predict function operates at the programme level, the context-specific failures your examiners will look for are outside its visibility.
QUESTION #7: Do you have a complete inventory of AI tools in operational use across your compliance function?
This is the question that produces the most uncomfortable answers – not because institutions are negligent, but because the adoption of accessible AI tools across compliance functions has outpaced governance programme visibility in almost every institution.
Compliance analysts use large language models to summarise regulatory guidance. Collections teams use AI-assisted drafting for customer communications. Risk teams use AI-generated analysis in regulatory submissions. None of this is malicious. All of it creates exposure.
The EU AI Act’s transparency and human oversight requirements apply to AI in operational use, not to AI that has been formally approved. The institution’s accountability for a regulatory submission extends to every element of that submission, regardless of which tool generated the first draft. A SAR narrative that cannot be validated against the underlying transaction data is not a compliant SAR, regardless of whether the tool that generated it appears on an approved technology list.
The shadow compliance AI estate in most institutions is larger than the formally governed estate. If you do not have a complete inventory – including unsanctioned tools – you do not know the full scope of your regulatory exposure.

THE COMMON THREAD: All seven questions share the same underlying structure: they ask whether the institution’s compliance governance architecture has kept pace with the granularity at which its AI actually operates. Per-entity not per-segment. Sentence-level not document-level. Decision-level not model-level. Context-specific not programme-level. The gap between those two levels of granularity is where examination findings originate.
Why These Questions Are the Right Ones to Ask Now?
The regulatory examination standard has moved from documentation to demonstration. From “do you have governance?” to “can you show, in production data, that your governance is operational?”
These seven questions are the operational expression of that standard. They are the questions a demanding examiner asks. They are the questions a technically informed board member should be asking. And they are the questions that, answered honestly, locate an institution precisely on the spectrum from compliance automation to compliance assurance.
The distance between where most institutions are and where the regulatory standard now sits is closable. But it requires building governance at the granularity AI actually operates – not the granularity the previous technology era required
Every institution accumulating the gap between automation and assurance is accumulating it silently, at the speed of its own AI deployment. The examination that finds it is not a future event. It is a programme already underway at every major prudential supervisor.
The compliance function is where the promise of AI-first banking is tested. Every governance decision made in model design, data architecture, and delivery pipeline becomes visible here – not in a technology review, but in a regulatory examination.

The architecture behind all seven answers:
- determination-level explainability,
- per-entity AML monitoring,
- sentence-level policy lineage,
- unstructured fair lending governance,
- report data lineage,
- context-specific resilience, and
- the shadow AI governance framework
is examined in depth in CIO Mandate Series Paper 4. Download: The CIO’s Guide to AI Compliance and Regulatory Governance