PAN Lab example
Danske Bank fraud scoring
Better detection but worse reimbursement — and a rule that moved it
A bank replaces a legacy fraud system - 40% detection, 99.5% false positives - with a real-time deep-learning engine claiming far better detection. Modeled on a documented rollout. But the same bank ranked worst at reimbursing scam victims, and what fixed that was not a better model: it was a regulator's rule. Watch the two levers this case keeps distinct - detection quality, held by the model and the analyst, and the justice of the disposition, held by a policy no classifier reaches.
Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.
What this models
This example runs on the Deep-learning fraud scoring with a separate reimbursement lever network: 5 components and 12 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.
Evidence base: 5 assumed · 2 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.
- assumed
Domain caveat carried onto the diagram (evidence dossier, honest gaps): there is no empirical study measuring analysts copying prior or peer alert dispositions in this domain. The peer pathway drawn here is therefore not peer imitation - it is a documented hand-off between distinct functions, the fraud analyst handing a case to the reimbursement function that decides the victim outcome - and the peer-governance lever is offered as governance of that seam, not as a remedy for disposition copying. The disposition-feedback dynamic the domain does have empirical grounding for is the retraining loop from the case record into the model, which is drawn on the record-to-model pathway and addressed by provenance labelling.
- assumed
The payment regulator is drawn as the review step because that is what the record documents - an external body whose published performance tables read this bank's reimbursement outcomes - and it is now wired in through that published record rather than left emitting a check it never receives anything for. It is deliberately not an internal supervisory tier: the record lists one frontline analyst class, and the decisive victim-outcome move was made by a rule the regulator wrote, an actor outside this organization's boundary.
- assumed
The score-to-analyst pathway is drawn at full strength from the measured legacy baseline this engine replaced: roughly 40 percent detection at 99.5 percent false positives, the domain's one measured figure rather than a vendor claim. The machine write to the case record is raised to a substantial level because scoring is real-time, under 300 milliseconds, so the score lands in the record without a person in between. The detection-side check stays empty because what exists is a vendor-published case study, not an independent audit. Workload runs above staffing - a heavy alert load worked by a single analyst class.
- baseline
This models the two-sided detection-vs-reimbursement pattern documented in the case file - not a reconstruction of the actual system. Its defining feature is that detection quality (the model + fraud analyst) and disposition justice (the reimbursement function) are different levers held by different actors: the same institution improved its in-line scoring and ranked worst among UK banks for reimbursing scam victims.
- baseline
The claimed detection improvement (-60% false positives, +50% detection, <300ms) is a vendor case study entered as a claimed magnitude, drawn against the legacy 40%-detection / 99.5%-false-positive baseline - which is the record's most credible datum precisely because a 99.5% false-positive rate is measured reality, not a marketing figure. The independent audit of the claim is drawn empty.
- assumed
The decisive datum is which lever moved the victim outcome: not a better model, but a rule. The regulator's mandatory-reimbursement regime raised sector reimbursement from ~two-thirds to ~89% - drawn as the latent oversight check on the reimbursement function, held by an actor outside the bank. This is why 'improve the model' does not reach the victim outcome: the reimbursement gap was never a detection problem, and no classifier improvement substitutes for the rule that closed it.
- assumed
No customer or scam-victim outcome is modeled here. This Lab reads institutional propagation only, and customers and victims are boundary-only. The detection figures, the reimbursement ranking, and the rule-change effect live in the case file, and are never computed from anything in this diagram.
What this example does not show
- No customer or scam-victim outcome is modeled. The Lab reads institutional propagation only; customers and victims are boundary-only, and the detection figures, the reimbursement ranking, and the rule-change effect live in the case file, never computed on this diagram.
- The -60%-false-positives / +50%-detection figures are a vendor case study entered as claimed magnitudes; the credible datum is the measured legacy 99.5% false-positive baseline, and the victim-reimbursement outcome is a regulator-published ranking moved by a rule, none of it computed by the diagram.
Sources and evidence
What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.
A Nordic bank's rules-based legacy fraud system ran at roughly 40 percent detection with a 99.5 percent false-positive rate — a measured pre-machine-learning baseline whose badness is the most credible datum in the record, since a 99.5 percent false-positive rate is not a marketing claim. The vendor-published rollout of a deep-learning engine scoring transactions in real time (under 300 milliseconds) claims false positives cut by about 60 percent and true-positive detection raised by about 50 percent; those figures are an organization-named, trade-press-covered vendor case study, entered here as claimed magnitudes against that legacy baseline because they were not independently audited.
empirical- Vendor Teradata (2017). Danske Bank Fights Fraud with Deep Learning and AI (case study EB9821). https://assets.teradata.com/resourceCenter/downloads/CaseStudies/CaseStudy_EB9821_Danske_Bank_Saves_Millions_Fighting_Fraud_With_Deep_Learning_and_AI.pdf
- Trade press Groenfeldt, T. (2017, October 30). Danske Bank Uses Tech To Prevent Digital Fraud. Forbes https://www.forbes.com/sites/tomgroenfeldt/2017/10/30/danske-bank-uses-tech-to-prevent-digital-fraud/
The same institution that improved its in-line fraud scoring later ranked worst among UK banks for reimbursing victims of authorized-push-payment scams in the regulator's bank-by-bank performance data — better detection and worse victim-outcome performance coexisting in one organization. And it was a rule, not a model, that moved the institutional behavior: the regulator's mandatory-reimbursement regime raised sector reimbursement from roughly two-thirds to about 89 percent, demonstrating that detection quality and the justice of the disposition are different levers held by different actors, and that the victim-outcome lever is a regulatory rule rather than a better classifier.
empirical- Government UK Payment Systems Regulator (2023-2025). APP fraud performance data / APP scams performance reports. https://www.psr.org.uk/information-for-consumers/app-fraud-performance-data/
Where this connects
Institutional pressures in this domain
- Workload surge — Demand outruns staffing; per-case attention shrinks and review becomes triage.
- Reviewer bottleneck — One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.
- Vendor opacity — The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
- Data & policy drift — The world, the intake process, and the rules change under a system trained on how things used to be — two mechanisms with different remedies: the statistical properties of what the system processes move (concept drift), or the mixture of inputs arriving in deployment differs from the mixture it was trained on (covariate shift).
- Compliance over substance — Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.
All of them in context on the Security operations & fraud detection domain page.
Levers available here and the patterns behind them
- Gate vendor updates — Vendor quality gate
- Mark AI-written records — Provenance labeling
- Peer sharing rules — Peer-edge governance
- Review on schedule — Oversight cadence & retrospectives
- Check with a second model — Cross-model verification
- Escalate checks — State-feedback vigilance
- Review the riskiest first — Risk-tiered oversight
- Store less data — Data minimization