Skip to content
Open to board advisory and board seats — 2H 2026, then CY 2027–2028.
See details →
AI

A Convincing Voice Is Not an Authenticated One

A synthetic voice, a cloned face, and a matching caller-ID are recognition, not authentication — and AI just made recognition cheap to forge. The control that still holds isn't a better deepfake detector; it's moving the trusted event off the content and onto a channel you control: callback on a number of record, a pre-shared challenge, dual authorization on money movement.

By Michael YorkJuly 20, 2026 9 min read 2,134 words All AITable of contents

A deepfake does not defeat your controls. A procedure that accepts a familiar voice as proof of identity defeats them for it. Move the trust off the content and onto the channel.

In a widely reported 2024 case, a finance employee at the engineering firm Arup joined a video call with people who looked and sounded exactly like the company's CFO and several colleagues. Every face on the call was synthetic. Convinced by what he saw and heard, he authorized a series of transfers later reported at around 25 million dollars. Nobody hacked a system. The controls did what they were told. A human recognized his coworkers and acted on the recognition, and the recognition was manufactured.

Treat that as a pattern, not a headline. FinCEN warned in a widely reported 2024 alert about deepfake media used to open accounts and move money, and the FBI has repeatedly flagged voice cloning and synthetic video in business fraud. As reported, the ingredients keep repeating: a cloned voice on a phone call, a synthetic face on a video, a caller-ID that matches, an email thread that reads exactly right. None of it is exotic anymore. It is a paid feature.

I run information security and platform teams in a regulated fintech that serves more than 1,500 financial institutions, and the question I hear in nearly every conversation about this is the same one the vendors are racing to sell an answer to: can we detect the deepfake? Can we buy a model that flags the synthetic voice, the generated face, the cloned video, before someone acts on it. It is the wrong question. It stakes your defense on winning a detection arms race against a technology whose entire job is to get better at not being detected. The question that actually protects you is narrower and unglamorous: where in our processes does recognizing someone stand in for proving who they are — and can we remove it?

Detection is an arms race, not a control

I have written before that AI's real effect on attackers is that it raises the floor: it collapses the cost of a fluent, personalized lure toward zero and hands the least-skilled attacker a capability they did not have. Synthetic media is that same shift applied to identity signals. The lure, the voice, and the face used to be three separate kinds of work, each with its own cost and skill. AI collapsed all three at once and dropped them to the price of a subscription. What rose is not one attack. It is the credibility of every channel a human uses to decide they are talking to who they think they are.

The instinct is to meet a detection problem with a detection tool, and there is a whole market forming to sell you one. Be careful. Deepfake detection is probabilistic, and it is adversarial: every detector is a target, and the generator is trained to beat exactly that kind of scrutiny. Liveness checks — blink, turn your head, read these digits — are already being defeated by injection attacks that feed a synthetic video stream straight into the camera path, bypassing the lens the check assumes exists. A detector that is 95 percent accurate sounds excellent until you remember the attacker retries for free and only needs to win once. You do not build a money-movement control on a coin flip the adversary gets to reweight.

Detection has a place — as telemetry, as a signal that raises friction, as one input among several. It does not have a place as the thing standing between a request and a wire. Anything you would stake a transfer on has to hold when the fake is perfect, because eventually it will be.

Recognition is not authentication

Here is the reframe I want to make stick, because it is the whole piece. A familiar voice, a recognizable face, a caller-ID you know, an email thread in the right tone — these are all recognition. Recognition is a probabilistic human judgment: this sounds like, looks like, reads like the person I expect. Authentication is something else entirely. It is proving identity by presenting a factor bound to the real person — a key, a passkey on an enrolled device, a callback to a number that person controls, a secret only they hold. We spent decades learning not to trust recognition for machines. We never finished the job for people.

I have made the machine version of this argument before: your AI agents can authenticate but they cannot prove they are authorized, and the credential quietly becomes the entire identity. The human version is the mirror image. A person on a call can be recognized but not authenticated, and for most of banking history we let recognition quietly become the entire identity — because forging a voice or a face convincingly was hard enough to treat as proof. AI removed that friction. The moment recognition is cheap to forge, every process that accepted "I recognized them" as authentication is now a process with no authentication in it at all.

So the audit is simple to describe and uncomfortable to run. Walk every path where money moves, access is granted, credentials are reset, or a payee is changed, and mark each point where the deciding factor is that someone recognized a voice, a face, a name, or a number. Every one of those marks is a control that AI just voided. You are not hunting for the deepfake. You are hunting for the places you were counting on nobody being able to make one.

Move the trust to the channel, not the content

Once you stop trusting the content — the voice, the face, the caller-ID, the thread — the question becomes what you can trust, and the answer is a channel you control the endpoint of. The content of a call is forgeable. A call you place to a number of record is not, because the attacker does not answer at the real CFO's phone. This is the oldest idea in fraud prevention wearing new urgency: out-of-band verification, initiated by you, to an endpoint enrolled before the request existed.

These are the controls I would put between recognition and action, and none of them require a new vendor:

  • Callback on a number of record, never a number you were given. The verification value of a callback is entirely in who controls the endpoint. Call the CFO on the number in your directory, not the one in the email or on the caller-ID. If the request is legitimate, you lose thirty seconds. If it is not, the person who answers has no idea what you are talking about, and that is the whole point.
  • A pre-shared challenge, and a duress word. A code phrase agreed in advance, out of band, that a synthetic caller cannot know because it was never spoken over a forgeable channel. Pair it with a separate word that means "I am being coerced," because deepfakes are not the only way a real voice ends up making a fraudulent request.
  • Dual authorization on money movement above a threshold. Two people, two independent approvals, for any transfer past a line you set. This is the human blast-radius control: one fooled employee is no longer sufficient to move the money, and an attacker now has to compromise two people through two channels instead of convincing one on a video call.
  • Payee and bank-detail changes verified out of band, every time. The highest-yield fraud is not a dramatic transfer, it is a quiet edit to where a legitimate, recurring payment lands. Any change to a payee's account details triggers an independent callback to a known contact before it takes effect — no exceptions for urgency, because urgency is the lever.
  • Step-up authentication bound to an enrolled device, not to a voice. When you need a higher-assurance factor, reach for a passkey or a push to a device the person enrolled, which proves possession of a thing the real person holds. A voiceprint proves only that something produced the right sound, and something now can.

Notice the common shape. Every one of these moves the trusted event off the content of the interaction — which the attacker authored — and onto a channel whose endpoint you established in advance. That is the same boundary I build for automated systems: the point where a request becomes an action is the point you instrument and gate, not the point where you decide the request sounded sincere. Sincerity is now generated. The channel is what is left.

Your call center is the perimeter now

For a financial institution, the softest place this lands is not the wire room. It is the help desk and the account-recovery flow — the people whose entire job is to be helpful to someone who says they are locked out. That agent is measured on handle time and customer satisfaction, and they are now being called by a voice that sounds exactly like an account holder, in distress, with a plausible story and half the right answers. Social engineering against the help desk is an old technique. AI just gave it a perfect voice and infinite patience.

So the control cannot live in the agent's judgment, because judgment is exactly what the deepfake targets. It has to live in the script — a procedure that never accepts recognition as a factor. Roughly the sequence I would enforce:

  1. Never authenticate on voice recognition, or on knowledge an attacker can research or phish — a name, a date of birth, the last four digits, a recent transaction. Treat all of it as public.
  2. Drive high-risk actions — password reset, MFA re-enrollment, payee change, contact-info change — to a step-up bound to an enrolled device, or a callback to a number of record, not to the call in progress.
  3. Give the agent an explicit, blameless path to slow down: a hold, a callback, a supervisor, with metrics that reward the friction instead of punishing the handle time.
  4. Log the verification method used, not just the outcome, so the pattern of how identity was actually proven is reviewable after the fact.

The same logic runs through onboarding. Synthetic-identity fraud and deepfake-defeated liveness are the account-opening version of the same problem: a generated face passing a check that assumed a real one. The answer is not a better face detector alone; it is binding the identity proof to signals harder to synthesize — a verified device, a bank-account link, a channel you control — and refusing to let a single passed liveness check be the whole gate.

A procedure gap is an examiner's question

When you serve more than 1,500 financial institutions, this stops being an internal fraud problem and becomes a design obligation. Every one of those institutions runs a call center and a payments desk, and every consumer they serve is reachable by a cloned voice. The verification procedures are not back-office hygiene; they are the control an examiner asks to see exercised. FFIEC authentication guidance and the GLBA Safeguards Rule already expect layered, risk-based identity assurance, and dual control on payments is table stakes in any serious fraud program. The US Treasury's Financial Services AI Risk Management Framework — 230 control objectives across seven domains, published in February 2026 — pushes the same way: demonstrate the control, do not just assert the policy. "We train our people to spot fakes" is a policy. "No single recognition event can move money, and here is the callback log that proves it" is a control.

What to change before the next call

You cannot buy your way out of this with a detector, and you cannot train your way out of it by asking people to be more suspicious of a perfect fake. You engineer the recognition out of the decisions that matter.

Map every path where money moves or access is granted and mark where recognition substitutes for authentication. Move the trusted event to a channel you control: callback on a number of record, a pre-shared challenge, dual authorization above a threshold. Rewrite the help-desk and recovery scripts so no agent can authenticate on a voice. And treat any deepfake detector you buy as telemetry that adds friction, never as the gate that stands between a request and a wire.

The deepfake is not the breach. The procedure that trusts the face is. Fix the procedure and the perfect fake becomes a strange phone call that goes nowhere.

If you have rewritten a verification procedure to assume the voice is fake — a callback policy, a help-desk script, a dual-control threshold — I want to hear where it created friction people actually routed around, because that is the part the guidance never covers. Tell me in the comments.

AIAI SecurityFraudDeepfakesFintech