PAN Lab example
ML anti-money-laundering as primary monitoring
Fewer alerts and more confirmed — but confirmed by whom?
HSBC replaced rules-based anti-money-laundering monitoring with Google Cloud's AML AI. It reports more confirmed suspicious activity from fewer alerts, figures no one has independently audited.
See more
AML AI is Google Cloud's machine-learning product for anti-money-laundering (AML) transaction monitoring. HSBC runs it as its primary monitoring system in key markets, in place of monitoring by fixed rules. It scores activity for risk and raises the alerts the bank's investigators work.
What HSBC reports
HSBC reports that AML AI found two to four times more confirmed suspicious activity, with roughly 60 percent fewer alerts. It also reports that the time to process data in large batches fell from weeks to days. The figures come from Google Cloud's press release of 21 June 2023, which launched the product, and from a case study by the analyst firm Celent.
These are the vendor's and the customer's own figures. No independent audit of them is on the public record. No peer-reviewed measurement of benefit across a real deployment, inside a named financial-crime operation, is public.
Why fewer alerts is first a workload figure
Money laundering is rare compared with the number of transactions a bank monitors. When the target is that rare, false alarms can make up most alerts, even from a very accurate system. Axelsson set out this effect, called the base-rate fallacy, in a 2000 study of intrusion detection, the detection of break-ins to computer systems.
The alert threshold is the risk score above which AML AI raises an alert. So moving it mostly changes how many alerts investigators must work. It changes much less how much laundering is caught. A cut of roughly 60 percent in alerts is first a claim about workload. It is a claim about accuracy only if something independent confirms it.
What more confirmed activity can mean
Card-fraud research by Dal Pozzolo and colleagues, published in 2018, found that only a small set of flagged transactions is ever checked against what really happened. Such models are retrained on the investigators' own decisions. Each decision becomes a training label, an answer the model learns from. The sources read for this case do not show how AML AI is retrained.
If AML AI learns that way, its alerts and its training labels can come to agree more over time without becoming more correct. A rising count of confirmed suspicious activity would then partly measure what the system taught its reviewers to confirm.
Who decides
HSBC's financial-crime function chose AML AI as its primary monitoring system. It sets the alert thresholds, staffs the investigations, and runs escalation. Google Cloud builds and tunes the model. It also published the deployment's performance figures, as HSBC reported them.
Financial regulators supervise the bank's anti-money-laundering work. No published independent evaluation of this model's performance exists. Supervision examines the bank's compliance function, not how precise the model is.
The bank's history
HSBC's prior anti-money-laundering enforcement history includes a 2012 deferred prosecution, in which prosecutors put charges on hold. So this deployment operates inside a compliance function that regulators have found wanting before. The cut in alerts that HSBC and Google Cloud advertise is exactly what a regulator worried about under-monitoring would question.
What the record supports
AML AI may help a great deal. The record supports a narrower reading. The benefit is claimed, not audited. The cut in alerts is a workload figure that a rare target can make look like accuracy. The confirmed count is tangled up with investigators' decisions on AML AI's own alerts.
Where to look
Watch the pathway from the case management record to AML AI, where recorded decisions become training labels. Then watch the two checks the record does not show: Independent audit and Check against real outcomes. They show what an independent record of what really happened would add.
What the available tools can and cannot address
A failure pathway is a link between two parts of the network where a mistake made by one part can be passed on to the other.
Explore (No Targets) sets no targets. Under Service Targets Only, the targets can be met. The cheapest combination costs 4 of this case's 12 budget units and uses two tools: Peer sharing rules and Escalate checks.
Under Service and Safety Targets and under All Governance Targets, the targets can also be met within the budget. Both levels ask you to close every failure pathway, among other targets. The cheapest combination costs 9 of the 12 units and uses four tools: Mark AI-written records, Peer sharing rules, Escalate checks, and Store less data.
Every combination of tools that fits the budget was checked. At each of those two levels, 24 of them meet the targets. Every one of the 24 includes Mark AI-written records, Escalate checks, and Store less data.
No tool offered here changes AML AI itself. The model is Google Cloud's product, so HSBC's hold over it is its terms with Google Cloud. Gate vendor updates stands for those terms.
Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.
What this models
This example runs on the ML-AML-class monitoring with a label-feedback loop network: 4 components and 12 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.
Evidence base: 5 assumed · 2 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.
Show all 7 assumptions
- assumed
No study measures anti-money-laundering analysts copying each other's decisions, or earlier decisions, on alerts. So the pathway from the alert investigators to the escalation and filing team is a hand-off between two teams, not staff copying one another's decisions. An investigator escalates a case to the team that confirms and files it. Peer sharing rules acts on that hand-off here, not on copying. Card-fraud research does document one loop: models retrained on investigators' recorded decisions. The network draws that loop from the case management record to AML AI. Mark AI-written records acts on that loop.
- assumed
The research model this example is built from has two groups of staff: frontline alert investigators, and a financial-crime team that handles escalation and filing. So the escalation and filing team is the review step, not a second frontline group, and it reads AML AI's alerts on escalated cases. No part stands for an independent auditor, because the record says no independent audit of these figures exists. The audit appears instead as a check that could be added, named Independent audit.
- assumed
The network assumes AML AI raises many more alerts than there are true cases, whatever the reported cut, because money laundering is rare among the transactions monitored. It assumes investigators check few flagged transactions against a real outcome, and that checked outcomes arrive late. It treats retraining on recorded decisions as a major source of shared error, because most flags are never checked against what really happened. It treats one model across key markets as a source of repeated error, because AML AI replaced the rules as primary monitoring. It assumes more alert work than the investigators can fully handle.
- baseline
This example follows the pattern the case file documents: a machine-learning system as a bank's primary anti-money-laundering monitor. It is not a copy of HSBC's actual system. Two effects shape it. First, money laundering is rare, so false alarms dominate the alerts, and moving the alert threshold changes the workload more than what is caught. That shows on the pathway from AML AI to the investigators. Second, investigators' recorded decisions become training labels, on the pathway from the case management record to AML AI.
- assumed
The network assumes investigators' recorded decisions become AML AI's training labels. Card-fraud research found that only a small set of flagged transactions is ever checked, and that models are retrained on investigators' own labels. If so, a rise in confirmed suspicious activity partly measures what the system taught its reviewers to confirm. The model and its labels can come to agree more without becoming more correct. That is why the check named Check against real outcomes matters most here. The sources read for this case do not show how AML AI is retrained.
- baseline
In the record for fraud and anti-money-laundering tools, every benefit figure measured across a real deployment comes from a vendor and its customer, with no independent audit. No peer-reviewed measurement inside a named financial-crime operation is public. The network shows that gap as the check named Independent audit. The reported cut in alerts is exactly what a regulator worried about under-monitoring would question. From the inside, a real saving of work and a rise in missed cases look the same.
- assumed
This example does not model any customer or enforcement outcome. It shows how errors move among AML AI, the bank's staff, and the case management record. Account holders and anyone flagged are outside the network. The reported figures, the point about rare targets, and HSBC's regulatory history come from the case file and its sources. Nothing in this example computes them. The benefit figures are entered as claims, not audited results.
What this example does not show
Show all 3 limitations
- HSBC's reported figures, two to four times more confirmed suspicious activity and roughly 60 percent fewer alerts, come from HSBC and Google Cloud themselves. No independent audit of them is on the public record. This example enters them as claims, not measurements. The check named Independent audit stands for the audit the record does not show.
- This example does not model any customer or enforcement outcome. It shows how errors move among AML AI, the bank's staff, and the case management record. Account holders and anyone flagged are outside the network. The reported figures, the point about rare targets, and HSBC's regulatory history come from the case file and its sources. Nothing in this example computes them.
- This example does not measure how well AML AI detects money laundering. HSBC's figures are self-reports, not audited results. The example shows the independent audit and the check against real outcomes as checks that could be added, because the record describes neither.
Sources and evidence
What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.
A global bank replaced rules-based transaction monitoring with a cloud vendor's machine-learning anti-money-laundering product as its primary monitoring system in key markets, reporting two to four times more confirmed suspicious activity with roughly 60 percent fewer alerts. Every one of those numbers is a vendor-and-customer self-report with no independent audit — which is itself the honest structure of the domain, because a peer-reviewed deployment-scale benefit measurement inside a named financial-crime operation does not publicly exist, and the alert-volume reduction the vendor advertises is precisely the lever a regulator scrutinizing an under-monitoring risk would question.
empirical- Vendor Google Cloud (2023, June 21). Google Cloud Launches AI-Powered Anti Money Laundering Product for Financial Institutions (with HSBC-reported results). https://www.googlecloudpresscorner.com/2023-06-21-Google-Cloud-Launches-AI-Powered-Anti-Money-Laundering-Product-for-Financial-Institutions
Two structural dynamics govern fraud and financial-crime detection. Under extreme base rates, detection precision is dominated by the false-alarm rate rather than by accuracy, so at realistic prevalence a threshold change moves the burden of alerts rather than the truth of them (the base-rate fallacy). And the labels the model learns from are the investigators' own dispositions: only a small set of flagged transactions is ever verified, and models are retrained on the analysts' calls, so a rise in 'confirmed' activity is partly a measure of what the system taught its reviewers to confirm rather than an independent ground truth (the label-feedback loop).
empirical- Academic Axelsson, S. (2000). The Base-Rate Fallacy and the Difficulty of Intrusion Detection. ACM Transactions on Information and System Security, 3(3), 186-205. https://doi.org/10.1145/357830.357849 https://dl.acm.org/doi/10.1145/357830.357849
- Peer-reviewed Dal Pozzolo, A., Boracchi, G., Caelen, O., Alippi, C., & Bontempi, G. (2018). Credit Card Fraud Detection: A Realistic Modeling and a Novel Learning Strategy. IEEE Transactions on Neural Networks and Learning Systems, 29(8), 3784-3797. https://doi.org/10.1109/TNNLS.2017.2736643 https://dalpozz.github.io/static/pdf/TNNLS_2017.pdf
Where this connects
Institutional pressures in this domain
- Workload surge — Demand outruns staffing; per-case attention shrinks and review becomes triage.
- Reviewer bottleneck — One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.
- Vendor opacity — The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
- Data & policy drift — The world, the intake process, and the rules change under a system trained on how things used to be — two mechanisms with different remedies: the statistical properties of what the system processes move (concept drift), or the mixture of inputs arriving in deployment differs from the mixture it was trained on (covariate shift).
- Compliance over substance — Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.
All of them in context on the Security operations & fraud detection domain page.
Levers available here and the patterns behind them
- Gate vendor updates — Vendor quality gate
- Mark AI-written records — Provenance labeling
- Assign a challenger — Structured dissent
- Peer sharing rules — Peer-edge governance
- Review on schedule — Oversight cadence & retrospectives
- Check with a second model — Cross-model verification
- Escalate checks — State-feedback vigilance
- Review the riskiest first — Risk-tiered oversight
- Store less data — Data minimization