Three invoices, one supplier, and an instruction nobody can verify by looking
A finance director is about to approve 822,990 US dollars in one click. The question this page is about is not whether they are allowed to. It is what the screen owes them before they do — and whether it owes them the same thing when the instruction arrived in a voice and a face that were generated.
Before the screen below matters, one assumption has to go. Corporate payment controls were built around a person who can tell whether an instruction is genuine. That person no longer exists.
A lookalike domain carrying a pixel-accurate copy of a real financial platform. Deposits work. Withdrawals never do. The interface is not the expensive part of this attack, it is the free part — and the harder the original works at looking trustworthy, the better the copy looks. What cannot be copied is the account the money lands in.
In 2024 a finance employee at the engineering firm Arup joined a video call with colleagues he recognised, including the group's chief financial officer. Every other participant was synthetic. Fifteen transfers followed, totalling around 25 million US dollars. Sourced below.
A cloned voice, assembled from public recordings, asks the account holder to read out the code that has just arrived by text. The code is the defect: a one-time password is a shared secret a person can say out loud, which makes it something that can be talked out of them.
Three different senses, one conclusion: recognition has stopped being evidence. Everything below assumes the person at the screen has already been convinced. The controls that survive that assumption are the ones that never asked whether the instruction looked genuine — they asked what shape the money was moving in.
Three invoices from one construction supplier, sitting in the queue on the same value date. The bank applies two thresholds: above 3,000 USD a payment attracts elevated authentication, above 500,000 USD it attracts enhanced authentication. Change any amount, change a payee, add a line — the chain below moves with it.
| Invoice | Payee | Amount (USD) | Alone | In this batch |
|---|
tier(batch) = max( tier(sum of lines), max over lines of tier(line) )The same three invoices, submitted separately and submitted as one batch. The comparison people expect is fewer approvals. The comparison that is actually true is fewer submissions, and a higher bar on the smallest line.
The video call was flawless. Fifteen transfers were not. Split an instruction into pieces that each sit under a per-transaction threshold and every piece clears — unless the threshold is also computed across a window.
It is the expression from the batch above with one word changed:
tier(batch) = max( tier(sum over lines), max over lines of tier(line) )tier(window) = max( tier(sum over window), tier(this transfer) )Every attack on this page ends in the same place: money reaching a beneficiary the payer did not intend. Two live schemes check the name on the receiving account against the name the payer typed — on the rails they cover.
Batching is a decision made by the payer that lands on the payee's ledger. One credit of 822,990 against three open invoices is an unallocatable payment — which is a worse outcome for the supplier than three separate credits would have been. ISO 20022 structured remittance information is what makes the trade even, and the failure state below is the one worth designing for.
remt.001 for the advice itself and remt.002 for advising where the advice can be
found. This removes the length limit. Its cost is that the money and the explanation now travel on separate
paths and may not arrive together, which turns a formatting problem into a timing problem.The friction in corporate banking is usually described as too many fields. That is the symptom. The defect is that the credential is bound to a device rather than to a role, in a process that requires six different people.
Three of these — company identifier, company number, banking number — are lookup keys, not secrets. They are asked for at every login because the system resolves who you are and which organisation you act for in the same transaction.
Separating identity from intent is what lets biometrics finally do useful work here. A fingerprint is a poor answer to which company do you act for and an excellent answer to is this the payment you meant. Signing over the payload rather than logging into a session is the difference between what-you-see-is-what-you-sign and a session someone else can steer.
A six-gate chain on device-bound credentials has no legitimate path for absence. Somebody is travelling, somebody has a new phone, somebody is on leave at quarter end and the payment run cannot wait. What happens next is not that the payment waits. What happens is that a credential gets shared, and the system records six distinct approvals that were not made by six distinct people.
That is worse than a slow process. A slow process is visible. A control that has been quietly delegated produces a log that looks perfect and means nothing — and the log is the artefact everything downstream trusts. A design that adds a seventh gate to a chain people are already working around has made the record less true, not the bank safer.
Three published figures survive checking. They are stated here with enough provenance to be disagreed with.
McKinsey reports that the average onboarding process for a new corporate client can take up to 100 days, drawn from quantitative surveys and interviews with two dozen global banks.
McKinsey & Company, Winning corporate clients with great onboarding, 5 October 2022. Source. Read as an upper bound on an average, not a typical duration.
Panko's synthesis of field audits conducted between 1995 and 2004 found that 94% of the 88 spreadsheets examined contained at least one error, with a mean cell error rate of 5.2% across the 43 for which it was reported.
Raymond R. Panko, What We Know About Spreadsheet Errors. The figure is at least one error, not serious error, and the audited sample is small and old — which is exactly why the number is quoted at 94 rather than rounded to a friendlier 90.
Pol's peer-reviewed assessment puts a defensible indicative mid-point at about 0.05% of criminal funds intercepted, against private-sector compliance spending of roughly 300 billion US dollars a year and around 3 billion recovered — compliance costs many tens to hundreds of times greater than the amounts recovered.
Ronald F. Pol, Anti-money laundering: The world's least effective policy experiment? Together, we can fix it, Policy Design and Practice, vol. 3 no. 1, 2020. DOI. This is the figure that makes the credential argument above matter: the apparatus generating the friction is not, on this evidence, catching much.
In 2024 an employee in the Hong Kong finance function of the engineering firm Arup took part in a video conference with people he recognised, including the group's chief financial officer, and was instructed to move funds. Every other participant on the call was AI-generated. Fifteen transfers followed, reported at around 200 million Hong Kong dollars, roughly 25 million US dollars. Arup confirmed the incident publicly and stated that fake voices and images were used.
CNN Business, Arup revealed as victim of $25 million deepfake scam involving Hong Kong employee, 16 May 2024. Source. Read as a reported incident with figures given by the press and confirmed in outline by the firm, not as an audited forensic account. It is cited here for one reason: it is the clearest public case in which every human verification method available to the person at the screen returned a false positive at once.
An account name-checking service for UK domestic payments, run as an API-based peer-to-peer service with no central infrastructure, available to regulated payment service providers holding accounts addressable by sort code and account number. More than 300 organisations participate. The Payment Systems Regulator mandated an expansion of coverage in 2024.
Pay.UK, Confirmation of Payee. Source. Volume and participant counts are as published by the scheme operator.
Article 5c of the Instant Payments Regulation requires payment service providers to offer the payer a service verifying the payee to whom the payer intends to send a credit transfer. It must be provided at no cost to the payer, applies to standard as well as instant credit transfers, and returns its result before the payment is initiated.
European Central Bank, Instant Payments Regulation. Source. The obligation date given here is the euro-area application date for the verification requirement.
The technical standards under the second Payment Services Directive require, at Article 5, that where a payer initiates an electronic payment the authentication code is specific to the amount and the payee, and that the payer is made aware of both. This is the regulatory anchor under the credential argument on this page: a code that authenticates a person can be read aloud to somebody doing something else, and a code bound to a payment cannot.
Commission Delegated Regulation (EU) 2018/389 of 27 November 2017, regulatory technical standards for strong customer authentication, Article 5. Instrument. Stated honestly: the instrument is linked, but the article text could not be retrieved to quote directly, so the requirement is described here from secondary legal commentary rather than reproduced from the primary text. Treat the description as accurate in substance and unverified in wording.
These appeared in my own first draft. None of them survived a search for a primary source, so they are listed here rather than deleted. Removing an unsourced number is not the same as being honest about it — and a reader who has seen these figures elsewhere deserves to know where the trail stops.
Traceable only to vendor marketing material from companies selling alert-reduction software. No regulator publication or peer-reviewed study found stating this range. The incentive of the only available source is to make the number large.
Traces to vendor-commissioned surveys with undisclosed sampling. Not independently reproducible.
The closest figure I could find is 55%, in trade-press coverage rather than a primary study. The 90 appears to be a repeated rounding of something else.
Cut entirely, along with the distributed-ledger section it belonged to. Not because the claim is impossible, but because neither number traces to a primary source, and because the section argued the wrong thing: a payment run of this kind fails on counterparty risk, custody and legal structure long before it fails on seconds. Speed was never the binding constraint here.
The observations behind it come from real production work on institutional banking software under confidentiality. Nothing proprietary appears here, no client is identified, and the thresholds, payee, invoice numbers and routing labels are illustrative. That is a real limit on what you can verify, and it is the reason this is published as an independent study rather than a named-institution case study.
3,000 and 500,000 are chosen so the arithmetic is legible, not because any institution uses them. The
max() rule is the transferable part; the specific numbers are not.
No usability testing, no error-rate measurement, no comparison against an existing approval screen. The claim that showing constituent lines improves approval quality is an argument, not a finding.
Truncation is modelled at a common field length to make the failure visible. Real behaviour depends on the specific channel, scheme and correspondent chain, and would need to be measured per corridor rather than assumed.
The window rule raises the authentication tier; it does not stop the money. On the default settings the first transfer still leaves before the control engages. What the design buys is a ceiling on the loss and a routing decision the impersonation cannot follow, and it is worth being plain that those are different things from stopping the fraud.
Set the interval wider than the window and nothing accumulates: every transfer is judged alone again and the rule never fires. A longer window catches the slower sequence at the cost of holding ordinary business at a high tier for longer. That trade-off is the whole design decision and there is no setting that escapes it.
An account opened in a deliberately adjacent name returns close match, which is the same answer thousands of legitimate payments produce every day. A warning that fires on honest traffic is a warning people are trained by experience to clear. The check is worth having and it is not a gate.
Every scenario on this page ends in an account that somebody opened and somebody onboarded. Nothing on the paying side reaches that. This page argues about the payer's screen because that is the surface a designer controls, not because it is where the control belongs.
Detection is an arms race against a generator that improves faster than the detector, and losing it once is enough. Nothing here inspects a face, a voice or a page for signs of being fake. The argument is that the payment leaves a shape that is checkable with arithmetic even when the impersonation is perfect — and that if the arithmetic is the defence, it should be stated as the defence rather than added quietly behind a detector that will eventually be beaten.
Whether a supervisor accepts aggregate-level authentication at all is a legal question with a jurisdiction-specific answer. This page argues about what the interface must show once the rule exists; it does not argue that any particular regulator would permit the rule.