Today's sponsor

Reply to everything. Edit nothing.

Your inbox is full. Slack is piling up. Client messages need a response yesterday. Typing thoughtful replies to all of it takes hours you don't have.

Wispr Flow turns your voice into clean, professional text you can send the moment you stop talking. Speak like you would to a colleague — tangents and all — and get polished output. Emails, Slack, LinkedIn, WhatsApp, whatever's open.

89% of messages sent with zero edits. Used by teams at OpenAI, Vercel, and Clay. Works on Mac, Windows, and iPhone.

Each edition delivers one clear, evidence-backed idea you can use. This week: Humans cannot reliably detect AI cloned voices- SMB affordable defenses

A caller with your chief financial officer's (CFO's) voice asks the help desk to reset a multi-factor authentication (MFA) enrollment. The voice is right. The cadence is right. The urgency feels real. Nothing about the call sounds wrong, because nothing about it is designed to. Voice cloning has quietly removed the one authentication factor most small businesses never thought to question: the sound of someone they know.

The Research

The peer-reviewed evidence on this is not ambiguous, and it does not favor defenders.

A 2025 study in Scientific Reports tested listeners against ElevenLabs voice clones across more than 200 speakers and found people are "poorly equipped" to identify artificial intelligence (AI) generated voices, failing both at matching a clone to its source identity and at judging whether a voice was natural or synthetic (Barrington et al., 2025). Listeners routinely attributed cloned voices to the real speaker, which is exactly the failure mode an impersonation attack depends on (Barrington et al., 2025).

A separate 2025 study in PLoS One found the same thing with clones built from roughly four minutes of source audio using commercial instant-cloning software (Lavan et al., 2025). Listener sensitivity was not significantly above zero in two of three experiments, and participants leaned toward calling ambiguous voices human. Cloned voices were labeled "human" between 58% and 70% of the time — statistically indistinguishable from how often listeners correctly identified genuine human voices (Lavan et al., 2025).

Machines do no better. Foundational work presented at the Institute of Electrical and Electronics Engineers (IEEE) International Conference on Acoustics, Speech and Signal Processing showed that speech-synthesis and voice-conversion attacks drove the False Alarm Rate of two state-of-the-art automatic speaker verification (ASV) systems as high as 99.11%, using attack models trained on just 24 short utterances — roughly five to six seconds of target speech (Wu et al., 2015). The authors were blunt: without countermeasures, these systems are "extremely vulnerable to spoofing attacks" (Wu et al., 2015, Conclusions section).

Detection research has not closed the gap. A 2025 survey in Sensors found leading audio deepfake detection techniques still struggle against rapidly evolving synthesis methods (Zhang et al., 2025). A blended detection framework evaluated on the ASVspoof2021 benchmark still produced a residual Equal Error Rate against synthesis-based Logical Access attacks, which remain measurably harder to catch than simple replay attacks (Sharafudeen et al., 2024).

The SMB Reality

Large enterprises can respond to this with budget — liveness-testing vendors, real-time detection pipelines, person-of-interest models. Those options exist in the National Security Agency (NSA), Federal Bureau of Investigation (FBI), and Cybersecurity and Infrastructure Security Agency (CISA) joint guidance, and they are aimed at high-priority individuals and larger organizations that can absorb the integration work (National Security Agency et al., 2023).

A small or medium-sized business (SMB) has a different set of constraints. Five to six seconds of a leader's voice is enough for an attack, and any company with a podcast appearance, a webinar recording, or a voicemail greeting has already published that much. Meanwhile the people most likely to receive the call — a bookkeeper, an office manager, a one-person help desk — are the same people authorized to move money or reset credentials, often with no second approver in the loop. The attack surface is small, which is precisely why a single successful call does so much damage.

The Practical Fix

Because voice cannot be trusted as an authentication signal (Wu et al., 2015; Barrington et al., 2025; Lavan et al., 2025), the effective defenses are procedural. The joint NSA/FBI/CISA advisory makes this point directly: real-time verification procedures, training, and response planning cut risk substantially without detection infrastructure (National Security Agency et al., 2023).

Three controls do most of the work:

  1. Callback verification. Any phone request touching money, credentials, or account changes gets verified by calling back a number already on file — never a number offered during the call (National Security Agency et al., 2023).

  2. Two-person approval above a threshold. Wire transfers and payment-detail changes require a second approver, confirmed on a separate channel. That single step removes the failure point attackers manufacture with urgency and authority.

  3. Voice-independent MFA. Credential resets and MFA enrollment are never approved on a phone call alone. Route them through a secure portal, a single sign-on (SSO) push notification, or in-person confirmation (National Security Agency et al., 2023).

The advisory also recommends one-time personal identification numbers (PINs), known personal details, or biometric checks before executing sensitive financial communications, plus tabletop exercises focused on finance and the help desk (National Security Agency et al., 2023). Detection tooling can wait until these three are enforced without exceptions.

Close

The cheapest control here costs nothing but a written rule and the willingness to enforce it when the caller sounds like your boss and says the wire has to go out in ten minutes.

AI Governance platform for SMBs: https://raic.rhindoncyber.com

References

Barrington, S., Barua, R., Koorma, G., & Farid, H. (2025). People are poorly equipped to detect AI-powered voice clones. Scientific Reports, 15, Article 11004. https://doi.org/10.1038/s41598-025-94170-3

Lavan, N., Irvine, M., Rosi, V., & McGettigan, C. (2025). Voice clones sound realistic but not (yet) hyperrealistic. PLoS One, 20(9), Article e0332692. https://doi.org/10.1371/journal.pone.0332692

National Security Agency, Federal Bureau of Investigation, & Cybersecurity and Infrastructure Security Agency. (2023, September 12). Contextualizing deepfake threats to organizations (Cybersecurity Information Sheet U/OO/199197-23). https://media.defense.gov/2023/Sep/12/2003298925/-1/-1/0/CSI-DEEPFAKE-THREATS.PDF

Sharafudeen, M., Andrade, J., & S., V. K. (2024). A blended framework for audio spoof detection with sequential models and bags of auditory bites. Scientific Reports, 14, Article 20192. https://doi.org/10.1038/s41598-024-71026-w

Wu, Z., Khodabakhsh, A., Demiroglu, C., Yamagishi, J., Saito, D., Toda, T., & King, S. (2015). SAS: A speaker verification spoofing database containing diverse attacks. In 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) (pp. 4440–4444). IEEE. https://doi.org/10.1109/ICASSP.2015.7178810

Zhang, B., Cui, H., Nguyen, V., & Whitty, M. (2025). Audio deepfake detection: What has been achieved and what lies ahead. Sensors, 25(7), Article 1989. https://doi.org/10.3390/s25071989