Skip to content

PAN Lab example

ML anti-money-laundering as primary monitoring

Fewer alerts and more confirmed — but confirmed by whom?

A bank replaces rules-based monitoring with a machine-learning anti-money-laundering (AML) system: 2-4x more confirmed suspicious activity, 60% fewer alerts. Modeled on a vendor-reported deployment. Both numbers are self-reported, and the arithmetic complicates both: when the target is rare, fewer alerts is a workload claim, not an accuracy one - and the model retrains on the investigators' own dispositions, so 'more confirmed' can mean the system taught its reviewers to confirm. Watch what an independent ground truth would show.

Stylized model of a documented deploymentSecurity operations & fraud detection

Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.

What this models

This example runs on the ML-AML-class monitoring with a label-feedback loop network: 4 components and 12 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.

Evidence base: 5 assumed · 2 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.

  • assumed

    Domain caveat carried onto the diagram (evidence dossier, honest gaps): there is no empirical study measuring analysts copying prior or peer alert dispositions in this domain. The peer pathway drawn here is therefore not peer imitation - it is a documented hand-off between distinct functions, the frontline investigator escalating a case to the tier that confirms and files it - and the peer-governance lever is offered as governance of that seam, not as a remedy for disposition copying. The disposition-feedback dynamic the domain does have empirical grounding for is the retraining loop from the case record into the model, which is drawn on the record-to-model pathway and addressed by provenance labelling.

  • assumed

    PAN's model-org record for this deployment lists two user classes - a frontline alert investigator and a financial-crime escalation tier whose stage is case escalation and filing - so the filing tier is drawn as the review step rather than as a second frontline group, and it receives an escalated-alert pathway of its own. The separate independent-audit component the earlier draft carried has been removed: an independent audit is documented as absent here, so it belongs on the latent independent model check, not on a component. A component for something the record says does not exist is a decorative one.

  • assumed

    The alert stream is drawn at full strength because precision here is floored by extreme base rates, so the flow to investigators is a flood whatever the vendor alert-reduction figure says. The investigator correction pathway is drawn faint because only a small subset of flagged transactions is ever checked and verification lags. The label-feedback pathway is drawn at full strength, the strongest link on the board, because it is the domain defining contamination: the model retrains on the investigators own dispositions while most flags are never ground-truthed. The monoculture pathway is drawn at a substantial level because this model replaced the rules engine as primary monitoring, so nothing independent disagrees with it. Workload runs above staffing - a heavy alert demand against limited investigative capacity - a flood that is partially, not wholly, worked.

  • baseline

    This models the machine-learning anti-money-laundering-as-primary-monitoring pattern documented in the case file - not a reconstruction of the actual system. Its two structural dynamics are the domain's foundations: extreme base rates (drawn on the alert-stream pathway, where the false-alarm rate floors precision so threshold levers move burden not truth) and the label-feedback loop (drawn on the case-record-to-model pathway, where investigator dispositions become the model's training labels).

  • assumed

    The label-feedback loop is the defining contamination: only a small fraction of flagged transactions is ever independently verified, and the model retrains on the analysts' dispositions, so a reported rise in 'confirmed suspicious activity' partly measures what the system taught its reviewers to confirm. The model and its labels can drift into agreement without becoming more correct - which is why the ground-truth verification check (latent here) is the one that matters.

  • baseline

    The independence absence is drawn on the independent model check, empty at baseline: every deployment-scale benefit number in this domain is a vendor-and-customer self-report with no independent audit, and no peer-reviewed measurement inside a named financial-crime operation publicly exists. The advertised alert-volume reduction is precisely the lever a regulator worried about under-monitoring would scrutinize - a workload win and an under-monitoring risk look identical from the inside.

  • assumed

    No customer or enforcement outcome is modeled here. This Lab reads institutional propagation only, and account holders and flagged parties are boundary-only. The claimed magnitudes, the base-rate arithmetic, and the regulatory history live in the case file, and are never computed from anything in this diagram; the benefit numbers are entered as claimed, not audited.

What this example does not show

  • Every deployment-scale benefit figure here - the 2-4x confirmed-activity and 60%-alert-reduction numbers - is self-reported by the vendor and its customer, with no independent audit on the record. They are entered as CLAIMED magnitudes, never as measured ones, and the latent independent-read check is where that absence sits on the diagram.
  • No customer or enforcement outcome is modeled. The Lab reads institutional propagation only; account holders and flagged parties are boundary-only, and the claimed magnitudes, the base-rate arithmetic, and the regulatory history live in the case file, never computed on this diagram.
  • The 2-4x-more-confirmed and 60%-fewer-alerts figures are vendor-and-customer self-reports entered as claimed magnitudes, not audited results; the diagram draws the missing independent audit and ground-truth verification as latent checks, it does not compute detection performance.

Sources and evidence

What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.

  • A global bank replaced rules-based transaction monitoring with a cloud vendor's machine-learning anti-money-laundering product as its primary monitoring system in key markets, reporting two to four times more confirmed suspicious activity with roughly 60 percent fewer alerts. Every one of those numbers is a vendor-and-customer self-report with no independent audit — which is itself the honest structure of the domain, because a peer-reviewed deployment-scale benefit measurement inside a named financial-crime operation does not publicly exist, and the alert-volume reduction the vendor advertises is precisely the lever a regulator scrutinizing an under-monitoring risk would question.

    empirical
    • Vendor Google Cloud (2023, June 21). Google Cloud Launches AI-Powered Anti Money Laundering Product for Financial Institutions (with HSBC-reported results). https://www.googlecloudpresscorner.com/2023-06-21-Google-Cloud-Launches-AI-Powered-Anti-Money-Laundering-Product-for-Financial-Institutions
  • Two structural dynamics govern fraud and financial-crime detection. Under extreme base rates, detection precision is dominated by the false-alarm rate rather than by accuracy, so at realistic prevalence a threshold change moves the burden of alerts rather than the truth of them (the base-rate fallacy). And the labels the model learns from are the investigators' own dispositions: only a small set of flagged transactions is ever verified, and models are retrained on the analysts' calls, so a rise in 'confirmed' activity is partly a measure of what the system taught its reviewers to confirm rather than an independent ground truth (the label-feedback loop).

    empirical
    • Academic Axelsson, S. (2000). The Base-Rate Fallacy and the Difficulty of Intrusion Detection. ACM Transactions on Information and System Security, 3(3), 186-205. https://doi.org/10.1145/357830.357849 https://dl.acm.org/doi/10.1145/357830.357849
    • Peer-reviewed Dal Pozzolo, A., Boracchi, G., Caelen, O., Alippi, C., & Bontempi, G. (2018). Credit Card Fraud Detection: A Realistic Modeling and a Novel Learning Strategy. IEEE Transactions on Neural Networks and Learning Systems, 29(8), 3784-3797. https://doi.org/10.1109/TNNLS.2017.2736643 https://dalpozz.github.io/static/pdf/TNNLS_2017.pdf

Where this connects

Institutional pressures in this domain

  • Workload surge — Demand outruns staffing; per-case attention shrinks and review becomes triage.
  • Reviewer bottleneck — One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.
  • Vendor opacity — The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
  • Data & policy drift — The world, the intake process, and the rules change under a system trained on how things used to be — two mechanisms with different remedies: the statistical properties of what the system processes move (concept drift), or the mixture of inputs arriving in deployment differs from the mixture it was trained on (covariate shift).
  • Compliance over substance — Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.

All of them in context on the Security operations & fraud detection domain page.

Levers available here and the patterns behind them

Documented case histories