Skip to content

PAN Lab example

Danske Bank fraud scoring

Better detection but worse reimbursement — and a rule that moved it

Teradata's case study claims Danske Bank's fraud engine catches more fraud. Yet the bank later ranked worst among UK banks at reimbursing scam victims.

See more

Danske Bank's fraud engine is deep-learning software the bank rolled out with the analytics company Teradata, described in a 2017 case study. It scores each transaction for fraud in real time, in under 300 milliseconds, and flags suspect ones to fraud analysts. It replaced a rules-based system.

What the engine replaced

Danske Bank is a Nordic bank that also does business in the UK. The sources read for this case do not say whether the engine runs in its UK business.

A rules-based system flags transactions by fixed rules that people write. Deep learning is a kind of machine learning that finds patterns in large amounts of data. The bank's old rules-based fraud system caught about 40 percent of fraud, with a 99.5 percent false-positive rate. A false positive is a legitimate transaction flagged as fraud.

The case file calls that the most credible number in this field's record. Nobody markets a false-positive rate that bad, so it reads as measured reality. It is the baseline the new engine's gains are claimed against. The figures themselves come from Teradata's case study.

What the engine is claimed to do

Teradata's case study claims the engine cut false positives by about 60 percent and raised detection of real fraud by about 50 percent. It scores in under 300 milliseconds. The case study names the bank, and Forbes, a business magazine, covered the rollout independently in October 2017.

The figures are still a vendor's case study, not an independent audit. This case treats them as claims.

Who decides what

Detection belongs to the engine and the bank's fraud analysts. Once a customer has lost money to a scam, the bank's reimbursement policy decides whether they are paid back. A different part of the bank holds that policy. The UK Payment Systems Regulator sets rules above it.

So better detection and worse treatment of scam victims can exist side by side in one bank. They are different things, held by different people.

How the bank ranked on reimbursement

The UK Payment Systems Regulator publishes bank-by-bank data on authorized push payment scams. In these scams, customers are tricked into authorizing a payment themselves. In that data, Danske Bank later ranked worst among UK banks for reimbursing the victims.

The sources read for this case do not give the bank's rate. They do not say which year's data ranked it worst.

What moved the outcome for scam victims

A rule moved it, not a better model. The regulator's mandatory reimbursement regime took effect in October 2024. Before the rule, the sector reimbursed about 65 to 67 percent. Under it, the share rose to 89 percent. The record does not say whether that share counts cases or money.

A classifier is a model that sorts each case into a category. A better classifier did not raise reimbursement across the sector. The mandate did. The record does not show Danske Bank's own rate under the rule.

What this case asks

Watch the two sides this case keeps apart. The engine and the analysts hold detection quality. A policy that no model changes holds the fairness of the outcome for a scam victim.

The case asks you to measure detection and the outcome for scam victims separately. A real detection gain should not stand in for a reimbursement record it does not touch. Some harms here are moved only by rules held outside the bank.

What the available tools can and cannot address

A failure pathway is a link between two parts of the network where a mistake made by one part can be passed on to the other.

Explore (No Targets) sets no targets. Under Service Targets Only, the targets can be met. The cheapest combination uses two tools, Mark AI-written records and Store less data, and costs 5 of this case's 11 budget units.

Under Service and Safety Targets and under All Governance Targets, this case is not fully addressable with the available tools. Those levels ask you to close every failure pathway, among other targets. No tool offered here acts on the pathway named Analyst decisions on flagged transactions. All 2,143 combinations of tools that fit this case's budget were checked, and each leaves that pathway open.

This case also offers no tool that redesigns the engine, which Teradata built. No tool writes the reimbursement rule, which belongs to the regulator outside the bank.

Stylized model of a documented deploymentSecurity operations & fraud detection

Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.

What this models

This example runs on the Deep-learning fraud scoring with a separate reimbursement lever network: 5 components and 12 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.

Evidence base: 5 assumed · 2 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.

Show all 7 assumptions
  • assumed

    No study read for this case measures fraud analysts copying earlier decisions or colleagues' decisions on alerts. So the pathway between people here is not copying. It is a hand-off from the fraud analysts to the part of the bank that decides reimbursement. The case file documents that the two belong to different parts of the bank. The sources read do not describe how cases are handed between them. Peer sharing rules are offered as rules for that hand-off, not as a cure for copying. The feedback this field does document is retraining: fraud models are updated from investigators' past decisions. This example places that feedback on the pathway from the fraud case record to the engine. The engine also writes its own scores into that record, so retraining could reuse the engine's own past output. Marking AI-written records acts on that pathway, so the engine's own scores can be told apart.

  • assumed

    The UK Payment Systems Regulator is the reviewer in this example, because that is what the record documents. It publishes each bank's scam reimbursement results, so it reads this bank's reimbursement outcomes from outside. The network connects it to the fraud case record through that publication. It is not a supervisor inside the bank. The record shows no supervisory tier inside the bank, and does not divide the analysts into separate teams. The rule that changed the outcome for scam victims across the sector was written by the regulator, outside the bank.

  • assumed

    This example assumes the engine's scores bear heavily on the analysts' work. That rests on the old system's figures: about 40 percent of fraud caught, at a 99.5 percent false-positive rate. The case file calls those the most credible figures in this field's record, though they too come from Teradata's case study. Scoring runs in real time, in under 300 milliseconds. So the engine writes each score into the record with no person in between. The record contains no independent audit of the engine, only the vendor's published case study. This example assumes more alerts than the analysts can work through.

  • baseline

    This example follows the pattern the case file documents: better detection alongside worse treatment of scam victims. It is not a copy of Danske Bank's actual systems. Its defining feature is that different parts of the bank hold detection and reimbursement. The engine and the fraud analysts hold detection. The reimbursement function holds the outcome for scam victims. The same bank improved its real-time fraud scoring and ranked worst among UK banks for reimbursing scam victims.

  • baseline

    The claimed gains are about 60 percent fewer false positives, about 50 percent more fraud detected, and scoring in under 300 milliseconds. They come from a vendor's case study, and this example enters them as claims. It sets them against the old system's 40 percent detection and 99.5 percent false-positive rate. The case file calls that baseline the most credible figure in the record, because nobody markets a false-positive rate that high. The record contains no independent audit of the claims.

  • assumed

    What moved the outcome for scam victims was a rule, not a better model. The regulator's mandatory reimbursement regime raised reimbursement across the sector from about two-thirds to 89 percent. This example places that rule as a check held by the regulator on the reimbursement function, from outside the bank. That is why improving the model does not change the outcome for scam victims. The reimbursement gap was never a detection problem. A classifier is a model that sorts each case into a category. No better classifier substitutes for the rule.

  • assumed

    This example does not model any customer or scam victim outcome. It shows how errors move among the engine, the bank's staff, the fraud case record, and the regulator. Customers and victims are outside the network. The detection figures, the reimbursement ranking, and the effect of the rule come from the case file. Nothing in this diagram computes them.

What this example does not show

Show all 2 limitations
  • This example does not show what happened to any customer or scam victim. It shows how errors move among the engine, the bank's staff, the fraud case record, and the regulator. Customers and victims are outside the network. The detection figures, the reimbursement ranking, and the effect of the rule come from the case file. Nothing in the network computes them.
  • The engine's gains, about 60 percent fewer false positives and about 50 percent more fraud detected, come from Teradata's case study. They are claims, not audited results. The old system's figures come from the same case study, and the case file treats them as the most credible numbers here. The ranking and the rule's effect come from the regulator's published data. The network computes none of them.

Sources and evidence

What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.

  • A Nordic bank's rules-based legacy fraud system ran at roughly 40 percent detection with a 99.5 percent false-positive rate — a measured pre-machine-learning baseline whose badness is the most credible datum in the record, since a 99.5 percent false-positive rate is not a marketing claim. The vendor-published rollout of a deep-learning engine scoring transactions in real time (under 300 milliseconds) claims false positives cut by about 60 percent and true-positive detection raised by about 50 percent; those figures are an organization-named, trade-press-covered vendor case study, entered here as claimed magnitudes against that legacy baseline because they were not independently audited.

    empirical
    • Vendor Teradata (2017). Danske Bank Fights Fraud with Deep Learning and AI (case study EB9821). https://assets.teradata.com/resourceCenter/downloads/CaseStudies/CaseStudy_EB9821_Danske_Bank_Saves_Millions_Fighting_Fraud_With_Deep_Learning_and_AI.pdf
    • Trade press Groenfeldt, T. (2017, October 30). Danske Bank Uses Tech To Prevent Digital Fraud. Forbes https://www.forbes.com/sites/tomgroenfeldt/2017/10/30/danske-bank-uses-tech-to-prevent-digital-fraud/
  • The same institution that improved its in-line fraud scoring later ranked worst among UK banks for reimbursing victims of authorized-push-payment scams in the regulator's bank-by-bank performance data — better detection and worse victim-outcome performance coexisting in one organization. And it was a rule, not a model, that moved the institutional behavior: the regulator's mandatory-reimbursement regime raised sector reimbursement from roughly two-thirds to about 89 percent, demonstrating that detection quality and the justice of the disposition are different levers held by different actors, and that the victim-outcome lever is a regulatory rule rather than a better classifier.

    empirical
    • Government UK Payment Systems Regulator (2023-2025). APP fraud performance data / APP scams performance reports. https://www.psr.org.uk/information-for-consumers/app-fraud-performance-data/

Where this connects

Institutional pressures in this domain

  • Workload surge — Demand outruns staffing; per-case attention shrinks and review becomes triage.
  • Reviewer bottleneck — One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.
  • Vendor opacity — The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
  • Data & policy drift — The world, the intake process, and the rules change under a system trained on how things used to be — two mechanisms with different remedies: the statistical properties of what the system processes move (concept drift), or the mixture of inputs arriving in deployment differs from the mixture it was trained on (covariate shift).
  • Compliance over substance — Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.

All of them in context on the Security operations & fraud detection domain page.

Levers available here and the patterns behind them

Documented case histories