Here is what happened. On a Tuesday morning in early 2025, a Brinks Home security customer received a phone call from what sounded like a support agent. The voice was polite, unhurried, and knew the customer's name, address, and the model number of their alarm panel. The caller explained there had been a firmware update and requested the one-time passcode sent to the customer's phone. The customer read it out. Within minutes, the attacker had access to the Brinks Home network, and eventually, to the records of 4.9 million customers.
The voice on the phone was likely not human. Or rather, it might have been a human, a harvested recording, or a synthetic clone generated in real time by an AI model. That ambiguity is not a side detail. It is the entire story.
We are used to thinking about phishing as a text problem: suspicious emails, malicious links, fake login pages. But the ground has shifted. Mandiant's 2025 incident response data now shows that vishing — voice phishing — has overtaken email as the primary initial attack vector for enterprise breaches. CrowdStrike separately reports a 442% increase in vishing attacks. Microsoft tracks ShinyHunters, one of the most prolific vishing-driven intrusion groups, targeting more than 1,000 organizations and eventually exfiltrating more than 1.5 billion records. This is not an anomaly. It is a structural shift in how attackers get in.
Here is the uncomfortable part. The same week those reports were published, Google began rolling out a feature called "Let Google Call." It is built on the same technology that made Google Duplex famous in 2018: an AI voice agent that can call a local business, ask about hours, book a table, or check on an order. But there is one key difference. This agent introduces itself as an automated caller. It says, "I'm an automated system calling on behalf of a user," and it expects the business to engage.
No red team reviewed this. No telecom standard was designed for it. No regulator signed off on it. Google is not just deploying a product; it is systematically training millions of business owners to trust a non-human voice on the phone. And that training is precisely the behavior that vishing attackers depend on.
The Trust Stack Attackers and AI Agents Both Use
Let me break down the technical stack of a modern vishing attack, because it looks almost identical to a legitimate AI voice agent.
First, natural speech synthesis. Attackers no longer need a human with a convincing accent or a script. ElevenLabs, OpenAI, and open-source models can produce a voice that has no telltale robotic cadence. The clone can be built from three seconds of a public video on LinkedIn or a recorded customer service call.
Second, contextually coherent dialogue. The old vishing script was stiff and rehearsed: "I'm from IT. Your password expired. Verify now." That does not survive contact with a savvy employee. But a language model can adapt in real time. It senses hesitation. It detects when the target asks a question the script did not anticipate. It adjusts its tone to match the emotional state of the victim. In the field of voice engineering, this is called conversational state tracking, and in the era of LLMs, it got dramatically better.
Third, urgency and routine. Both legitimate agents and attackers rely on a specific rhetorical move: the call feels like a routine check, but it requires a specific action — a code, a password, a payment confirmation, a forwarding rule. The victim has been conditioned to treat ordinary requests as benign.
Fourth, the request for a specific operation. The attacker may ask the target to visit a URL, approve an MFA prompt, or read out a code. Sometimes they do not even need that. Once the attacker has access to the victim's voice and the victim has said the word "yes" on a clear-enough recording, they have a biometric artifact that can be replayed on a bank's voice-authentication system.
Google Duplex is not a miracle of new AI architecture. It is an integration of existing automatic speech recognition, text-to-speech, and conversational models into a single product. The innovation is not in the model. The innovation is social: Google decided that an AI agent can insert itself into a flow of human trust as a legitimate actor, and that the receiver of the call should accept it.
Missing from the public discussion is the observation that no one has solved the verifiable identity problem for AI voice agents. The telecom industry has a standard called STIR/SHAKEN, which digitally signs calls so that the caller ID cannot be spoofed, but it authenticates the carrier and the number, not the entity or the agent. A legitimate AI agent from Google and a malicious AI clone from a cyber criminal both call from numbers that pass STIR/SHAKEN if the attacker uses a VoIP provider that does not validate its customers. The declaration "I am an automated agent" is a text claim, not a protocol guarantee.
In 2017, I audited a token contract that had raised millions, and I found an integer overflow in distribution logic that would have let an early user mint more shares than the cap. The developers fixed it. But the deeper lesson stayed with me: markets and technology can both be fooled by confidence. That is why I am not reassured by a polite robotic voice saying it is a robot. I check the signature on the blockchain, not the face on the screen.
Here is the counter-intuitive point that does not appear in most coverage: honest transparency may be the most dangerous new attack vector. If AI agents are trained to say "I am automated," then a malicious agent can say the same sentence and be more convincing. The stranger asks, "Are you a real person?" And the AI voice geminates for a split second, then answers, "No, I am an automated system following a script. I cannot be manipulated. But to verify your account, I need the code sent to your phone." That is a perfect lie. The victim's mental model says: an honest robot would not lie. But the robot is not from Google. It is from ShinyHunters.
The social contract around phone calls has always been asymmetrical. The caller holds more information than the receiver. In the pre-AI era, a human voice carried a kind of IRL-proof: speakers have accents, intonation patterns, breathing, a finite amount of token budget for small talk. A cloned voice and a language model erase those cues. The receiver has almost zero observable evidence to distinguish between an authorized agent and an attacker.
The Unseen Costs: What Google Is Really Training
The most important thing to understand about "Let Google Call" is that the model of the AI agent is not being fine-tuned in any meaningful way. The system already works. What is being trained is the behavior of the human on the other end of the line. Every successful legitimate AI call — where the business makes a reservation, answers a question, changes a password — is a data point in a large, distributed conditioning experiment. The receiver learns: see a call from an unknown number, pick up, answer politely, comply when it asks for something routine.
Bank robber Willie Sutton famously said he robbed banks because that is where the money is. The modern version of Sutton's law is: attackers call because that is where trust lives. Businesses have firewalls, EDR detection, and phishing-aware employees looking at email. But the phone line has no SIEM, no sandbox, no reputation score. It is the softest perimeter in the modern enterprise.
There is an economic incentive for Google to push ahead despite the trust deficit. The data captured during these calls — hours of operation, live inventory, prices, service-level details — feeds directly into Google Maps, search listings, and the local search ecosystem. The conversation is the product. The declared identity is just a marketing veneer.
The commercial consequences ripple further. Cybereason and Cisco Talos have documented that vishing attacks are now the gateway to ransomware. Once the attacker has a Salesforce or Okta admin session, they can move laterally, extract data, and deploy encryption. The Brinks Home breach, which also implicated ADT and EY in connected incidents, is an illustrative case. The ransomware group ShinyHunters did not get in through a sophisticated zero-day exploit. They got in through a phone call and a user's impulse to comply.
The Contrarian View: Retail vs. Smart Money
The retail narrative says: "Google is protecting users with AI agents that clearly disclose their machine identity. Attackers are just criminals. We need better AI detection." That narrative misses the layering problem.
Smart money is not spending on AI that detects AI voices. It is spending on identity architecture — FIDO2, passkeys, hardware-bound authentication. Or it is buying insurance, which is where the vishing risk is being priced. The cyber insurance market is waking up. Claims data from 2024 shows that vishing and business email compromise are the leading causes of paid ransomware claims. Premiums will rise. Coverage will tighten. That is the point of maximum opportunity for companies that can prove they have strong inbound-call verification.
The security industry has an inherent incentive to amplify the fear. CrowdStrike sells EDR, not voice detection. Mandiant sells incident response, not prevention. Both publish the vishing numbers because they want to be seen as the authoritative source of the threat landscape. And that is fine. But it means that the economic signal from their reports is not just about the attack trend. It is about which vendors will benefit from the subsequent budget allocations.
The retail defense strategy focuses on employee training. I have trained community members for years, and I have distilled one line: trust is the only asset that survives the crash, and also the only asset that attackers can steal. Security awareness programs that teach phishing via email have improved detection rates for text-based attacks. But when an employee is on a call, with a voice in their ear, the cognitive load is higher. The attacker can interrupt like a human, rush the victim, apologize for the inconvenience, laugh at a joke. The training does not transfer.
The real answer is not better training. It is better procedures. High-risk actions should never be done over the phone without a second factor that is independent of the voice channel. If an internal admin wants to make a change to a payment system, an email to approve should still be required, or a hardware key. No one should be authorized to read out a password over the phone, even to a real person.
Three Waves of Industrial Impact
The first wave is already breaking on the enterprise security industry. If vishing has overtaken email, then security budgets must rotate toward voice-channel detection and response. This is not just about that rare, high-value call from a fraudster. It is about monitoring the entire call volume for social-engineering patterns. That is a new product category: call-aware security. Several startups are building it, embedding LLMs to rate the suspiciousness of a call in real time and alert the SOC.
The second wave is in the identity industry. Weak authentication — SMS codes, voice-based verification, or simple callback confirmation — is losing its credibility. The attack surface is the very voice channel that carries the authentication. A FIDO2 passkey is one answer, but it cannot work in every scenario. We still need high-assurance confirmation for customer support calls, for wire transfers, for password resets. The industry needs a caller-side credential that can be verified at the terminal: a cryptographic attestation embedded in the SIP header, or a signed JSON Web Token delivered over a side channel. STIR/SHAKEN is the foundation, but it is not agent-aware. Extending it to attest the software and the entity behind an automated caller is the next frontier.
The third wave is in the AI voice agent industry itself. Here is the paradox: if the social conditioning continues, the receiver's default will be to distrust all calls. That would be a very difficult landscape for legitimate AI agents, which have a hard enough time convincing a business to take a booking. The legitimate industry will need the same identity infrastructure as the security industry. And they will need it faster, or they will be the collateral damage of their own success.
As a founder in the copy trading space, I have seen the equivalent pattern in financial markets. Many projects burn their community by over-promising and under-verifying. They claim they are transparent but do not prove it on-chain. We walk away from greed; we stay for trust. The same logic applies to voice agents. If Google cannot demonstrate cryptographic identity for its AI agents, it is unintentionally feeding the very environment that attackers need.
The Red Team Gap
I could not find any evidence that Google red-teams "Let Google Call" against vishing-style attacks. The company has published details about the dialogue system, the policy layer that decides when it is safe to automate, and the ML fairness evaluation. But none of those address an adversarial call scenario, where the person on the other end is not a restaurant manager but an attacker trying to extract information from the AI or turn it into a malware-delivery vector.
The more likely path is that a security researcher, or a funded attacker, will use a legitimate AI agent as a weapon. Imagine a call from a known, trusted number — the one from your bank — but on the other end, instead of a human, is a vishing-trained bot that has been designed to mimic the bank's hiring practices. The victim's caller ID says the call is genuine. The voice is calm and professional. The first 30 seconds are about a routine issue. Then the bot asks: "To confirm your identity, could you press 1 for your account details?" Now the attacker has access to the account. This is not science fiction. The technical ingredients exist today.
The Way Forward: A Four-Party Trust Architecture
Trust will not be restored by a single company. It requires coordination across four layers.
First, the telecom infrastructure: STIR/SHAKEN must be extended to carry an agent attestation field. A legitimate AI agent should be required to include a cryptographic attestation that identifies the platform that generated the call and the entity that initiated it. This field should be sent to the terminating network and, ideally, displayed to the receiver as a "verified automated agent" badge.
Second, the AI platforms: Google, OpenAI, Microsoft, and Amazon need to agree on a standard for signing the output of their voice agents. Every call generated by an AI platform should have an immutable transcript hash that can be cross-checked on a central registry. If the caller claims to be an AI but cannot demonstrate a valid signature, the receiver's phone should flag it as an unverified call.
Third, the regulators: The FTC and FCC have begun asking questions about AI voice cloning. They should mandate that any AI-generated call include a clear, upfront disclosure in the first 10 seconds of the call, and that the disclosure be verifiable through a public directory. The disclosure also needs to survive if the call is forwarded.
Fourth, the receiver-side behavior: users and businesses must adopt a "zero-trust phone" policy. That means having a secondary channel — an app on the phone, a browser extension, or a desktop notification — to confirm that a call is expected and legitimate. This is more than training; it is a design pattern.
The Personal Lesson: Every Scar Teaches a Rule
After the 2022 Terra Luna collapse, I hosted daily town halls in Lagos, taking responsibility for losses in my community, and we implemented a community-voted risk protocol. The lesson I repeat to every new member is: every scar in the market teaches a new rule. But markets do not have a monopoly on scars.
The Brinks Home breach is a scar on the voice channel. It teaches us that a human cannot be the only layer of verification in a transaction. Trust is the only asset that survives the crash — but in the digital age, trust must be encoded in protocols, not just in hearts.
There is a useful analogy to the early days of email. Before DKIM and DMARC, phishing emails were indistinguishable from legitimate ones. The fix was not to ask users to be more careful; the fix was to authenticate the sender. The same evolution is now happening in voice. We need, in essence, a DKIM for the ear.
Until then, the prudent rule is to treat every inbound call that requests any action as a potential attack. Do not give a code over the phone. Do not confirm a password over the phone. Do not approve a login request over the phone. Instead, hang up, open the official app, and verify the request in a channel you initiated.
This is not paranoia. It is pattern recognition. The only way to survive in a world where hackers and AI agents share the same vocal chords is to assume that no voice is a guarantee, and to build trust that can be verified beyond audio.
The Takeaway: What I Am Watching Now
In the next six to twelve months, I am watching three signals. First, whether Google, Microsoft, or Apple introduces a public cryptographic attestation standard for AI voice calls. If they do, that is a signal that the trust architecture is being built. If they do not, the trust deficit will continue to grow, and the floodgates for vishing will widen.
Second, whether telcos begin to label AI calls as a distinct category, separate from human calls, at the network level. That would be a practical measure to let consumers know they are interacting with a bot. It could be an opt-in feature, but it would be a powerful signal that the industry is treating AI calls differently.
Third, whether the cyber insurance market begins to require evidence of voice-channel verification for policy renewals. If that happens, it will force businesses to adopt the new protocols, faster than any regulator could.
The future of voice is not about making AI sound more human. It is about making trust more machine-verifiable. The paradox of the Brinks Home breach is that it exposed the missing infrastructure. We should not wait for the next billion-record breach to build it.
I have been in this industry long enough to know that we cannot automate trust. But we can automate its verification. That is the only way to protect the flock — not just the profits. Our community has built a protocol that makes every trade transparent; the voice industry needs the same. Until then, trust nothing on a call that you cannot verify off the call.
At the end of the day, an AI agent is a tool. A human is a tool too. The question is not what they say, but whether we can verify who sent them. The phone line was the last frontier of analog trust in a digital world. The next two years will decide whether we surrender it to the attackers or rebuild it with the infrastructure it always deserved.