Four point nine million records. That is the tab Brinks Home paid for a single vishing breach in 2025, and the bill is still compounding. Within weeks, the same playbook hit ADT and EY, two more names that trusted their perimeter controls more than they trusted their own phone lines. Mandiant now classifies vishing as the top initial intrusion vector, overtaking email for the first time in a decade of tracking. CrowdStrike logged a 442% year-over-year spike in voice phishing. ShinyHunters, the group reportedly behind the Brinks entry, has targeted more than 1,000 organizations and amassed what Microsoft estimates to be more than 15 billion records.
This is not a random criminal wave. It is a migration. Attackers are not just changing tools; they are switching trust substrates. And they are doing it at the exact moment the tech industry is normalizing automated voices in everyday commercial life. The same week these reports crossed my desk, Google began pushing its “Let Google Call” feature, where an AI agent phones a small business, politely states that it is automated, and expects the human on the other end to engage with it. We are not just witnessing a security trend. We are being systematically trained to accept the weapon.
I have seen this pattern before. In 2017, I spent six weeks reverse-engineering early ERC-20 token implementations during the ICO frenzy. I found a reentrancy flaw in a fundraising contract that had already processed $4.2 million in ETH. When I posted the technical critique, the response was a furious debate about speed versus security standards. The industry chose speed. The industry paid for it in 2022, when the narrative behind algorithmic stablecoins collapsed exactly the way forensic analysis said it would. The story is never about the code alone. It is about what humans are conditioned to believe.

Today, the conditioning is happening in real time, on phone lines, at a scale Google alone can deliver. And the security industry is still looking for malware signatures while the real intrusion vector has become the human reflex to answer, respond, and comply.
The Training Operation
Let us be precise about what Google actually built. Duplex was first demonstrated at Google I/O in 2018, making a hair salon appointment with a voice so natural the receptionist had no idea she was talking to a machine. Seven years later, the core components — automatic speech recognition, text-to-speech synthesis, dialogue state tracking, real-time intent classification — have been heavily industrialized. “Let Google Call” is not an architectural breakthrough. It is an engineering integration, a combination of mature models into a productized deployment. The intelligence inside the agent has not suddenly leapt forward.
The real breakthrough is a product decision, not a model decision. The agent proactively calls a business, announces its own automation, and expects the merchant to interact with it. That is the first time a major platform has deployed an AI actor as an active participant in the commercial world, rather than a passive assistant waiting for commands. The agent carries intent, makes requests, negotiates around business hours, and completes a transaction on behalf of a user. It is a social role, not a tool.
That shift has a shadow side that Google’s product brief almost certainly did not include. Every time a legitimate AI agent completes a call, it teaches the person on the other end that automated voices are normal, routine, and safe to interact with. It conditions the merchant to follow instructions from a non-human actor. This is not a side effect. It is the mechanism. And it is the exact mechanism that vishing attackers have spent the past three years perfecting.
Vishing is not a crude scam. Modern vishing campaigns are built on four layers: a synthetic voice that sounds convincingly human; a script that references the correct internal systems, vendor names, and employee titles; a framing device that creates urgency or routine legitimacy; and a specific action request, often a password reset, a token approval, or a wire transfer. Remove the malicious intent, and you have just described Google’s automated calling agent. The legitimate AI and the attacker share the same trust stack. The only differentiator is intent, and intent is not visible to the receiver.
The Shared Trust Stack
Let me break the stack down, because this matters more than any single vulnerability disclosure.
The first layer is natural-sounding voice. Google has spent years making synthetic speech indistinguishable from human speech. Criminal outfits now use voice cloning models with a few seconds of audio to do the same. Detection tools that rely on robotic cadence or audio artifacts are chasing a ghost. The legitimate agent is indistinguishable by design, and so is the attack.
The second layer is contextual dialogue. The AI agent knows the business it is calling, the service it wants, and the appropriate questions to ask. Attackers scrape the same public data. They know your security vendor, your CRM, your CFO’s birthday, and the name of the internal helpdesk tool. Context is no longer a signal of legitimacy. It is a commodity.
The third layer is the framing device. Legitimate AI calls frame themselves as routine: “I’m calling to confirm your hours,” “I’m checking on the status of an order.” Vishing calls frame themselves as urgent: “I’m from IT, we detected a breach on your account, I need you to confirm a code.” Both exploit the same cognitive shortcut: when a situation feels familiar, the brain skips verification.
The fourth layer is the action request. The Google agent asks for a booking or a price update. The attacker asks for an MFA code or a password reset. The human response pattern is identical: listen, process, comply. The narrative is the attack surface. The story the caller tells, whether true or false, is what moves the human into action.
I used to think the market’s most important narratives were the ones on-chain. In 2020, while back-testing yield farming incentives across Uniswap and Compound, I concluded that yield is just liquidity rental. Projects rent user capital with token emissions, and the rental price is governance power. The same logic applies here. Acceptance is trust rental. Google is renting the collective trust of the telephone network and paying for it in behavioral conditioning. Attackers are withdrawing from the same account, and they are not paying anything.
The Half-Truth Disclosure
Here is the uncomfortable part. Google’s agent announces itself as automated. That sounds like transparency. But a malicious caller can say the same sentence. “I am an AI agent” is a claim, not a credential. Once society becomes accustomed to AI calling businesses, that statement stops being a disclosure and starts being a disguise. A vishing attacker can simply adopt the phrase as cover, and the receiver has no protocol-level way to verify whether the claim is true.
This is the identity vacuum at the heart of the problem. The telecommunications industry has STIR/SHAKEN, a framework that signs caller IDs at the carrier level. But STIR/SHAKEN authenticates the telephone number, not the entity using it, and it has no native concept of an AI agent’s identity. There is no cryptographic registry of authorized AI actors. There is no way for a merchant to query a canonical record proving that this specific AI voice belongs to Google, is acting on behalf of a Google user, and has permission to make the request it is making. Without that, the disclosure is just text. Beautiful, trust-building text that costs the attacker nothing to copy.
Based on my audit experience, I can tell you exactly what happens when identity is optional. The 2022 Terra collapse was not a technical failure. It was a narrative failure. The protocol claimed decentralization while the economic architecture centralized risk into a single oracle. Every market participant believed the claim because the narrative was comfortable. When the comfort broke, the peg broke. The same dynamic is now playing out in voice. We are being asked to trust a claim — “I am an automated agent” — without a verification mechanism. The human ear is a terrible oracle.
There is another layer of risk the product brief does not mention. Modern LLM-based voice agents can detect emotion, hesitation, and even adapt their persuasion style in real time. Those capabilities are not exclusive to legitimate vendors. Attackers can jailbreak a commercial agent, feed it a vishing script, and let it operate at scale with the same emotional intelligence. Red teaming for this scenario is almost nonexistent. No major platform has published a test regime that simulates an attacker inducing an agent to disclose user data or approve a fraudulent transaction. The capability is assumed safe until proven dangerous, which is precisely the opposite of how security engineering is supposed to work.
The Market Fallout
The industry impact is not speculative. It is already visible in the quarterly reports of every security vendor that tracks intrusion vectors. EDR, SIEM, and XDR products were built to detect malware and lateral movement, not to stop a social engineer who calls the helpdesk and asks for a password reset. Security awareness training, meanwhile, is still dominated by email phishing simulations. If the initial vector has moved to voice, the training content must move too.
MFA is now a liability as much as a defense. An attacker who has stolen a password only needs to convince the user to approve a push notification or read out a numeric code. Voice-based MFA, especially when the user is already primed to trust an automated caller, becomes the final step in the attack rather than a barrier. The push toward FIDO2 and passkeys is real, but the transitional period is long, and in that gap, the phone is the most dangerous device in the enterprise.
The telecom layer is also behind. Operators have implemented STIR/SHAKEN unevenly, and AI-agent calls do not fit cleanly into the framework. There is no regulatory classification for “AI-originated call,” no mandatory labeling standard, and no enforcement mechanism for fake AI disclosures. China has imposed labeling requirements on AI outbound calls, but the US and EU have not caught up. In a regulatory vacuum, the attacker always moves faster than the compliance team.
Network insurers are quietly repricing. ShinyHunters’ campaign against Brinks, ADT, and EY is a systemic risk event, not an isolated incident. Claims involving social engineering and human error are traditionally treated as operational risk, and they have been systematically underpriced. That will change. Insurance will eventually require verifiable call identity as a condition of coverage, which will create a market for trust infrastructure. If you want a forward-looking signal, watch the policy language, not the press releases.
The Contrarian View
The conventional take is that consumers will reject AI calls when they feel creepy. I believe the opposite is more dangerous. The data tells us that 64% of US consumers distrust major platforms, and only 13% fully trust AI. But distrust does not produce silence. It produces selective acceptance. People still answer the phone when the caller ID looks plausible. They still transfer money when the voice sounds like their CEO. Selective acceptance is precisely what vishing needs. A targeted attack does not need to convince everyone. It needs to convince one privileged user at the exact moment their guard is down. A little trust is more dangerous than none.
This is where the hunt for alpha in the noise of the herd begins. The market herd is chasing the next deepfake detector, the next audio forensics tool, the next ML anomaly score. But the structural alpha is in trust repair, not trust detection. The product that matters is a verifiable identity layer for voice: a cryptographic signature that binds an AI agent to its operating entity, a public registry that merchants can query, and a consumer-visible indicator that a call is genuinely authorized. In the same way DKIM and DMARC gave email a mechanism to separate legitimate senders from impersonators, voice needs a handshake for agents. Whoever builds that handshake owns the next decade of communications security.
The story behind the token, not just the ticker, is the story of who underwrites trust. The ticker here is the phone number. The token is the identity proof behind the voice. Google has the infrastructure and the user base to define this standard, but transparency alone does not create trust. A competitor that moves faster on protocol-level identity could flip Google’s structural advantage into a weakness. The market is wide open.
I have seen this movie before in crypto. Tether dominated more than 70% of the stablecoin market for years without a truly independent audit, and the industry collectively agreed not to look too closely. We called it a liquidity tool. We called it convenient. We pretended that a reserve claim was the same as a reserve fact. The peg held until it was tested, and when it was tested, the trust expense came due with interest. Voice authentication is heading the same way. The industry will accept “I am an AI” as a sufficient credential, the way it accepted Tether’s word as sufficient collateral, until an attacker proves the claim is worthless at scale.
The Takeaway
The fix will not be another firewall. It will not be a model that detects synthesized speech. The human ear is already outmatched, and the human reflex to comply is already being trained by the very companies we should trust most. The fix is a change in default: from “trust every voice” to “verify every actor.” That requires telco-grade authentication, platform-level signing, and a new category of call-object security that treats every incoming voice interaction as unverified until proven otherwise.
If we keep conditioning humans to say “yes” to any politely robotic voice, the next 15-billion-record breach is just a scheduled transaction. The question is not whether AI can sound human. It clearly can. The question is whether a human can prove they are not being played by a machine pretending to be an AI. Build the handshake. The narrative will follow.