JournalResearch

Field guide / 6

At least 26.9% of unwanted calls in one honeypot opened with a machine voice

A new 10,987-call study finds that replayed recordings, fresh synthetic speech, and silent machine connections demand different evidence, not one confident detector score.

Sep 11, 20266By ISH Team
At least 26.9% of unwanted calls in one honeypot opened with a machine voice
Advertisement

At least 26.9% of unwanted calls in one honeypot opened with a machine voice

An unfamiliar number calls. The voice sounds natural, pauses in plausible places, and knows your name. None of that tells you who is speaking. It may not even tell you whether anyone is speaking live.

Some automated callers replay a fixed recording. Others synthesize fresh speech. A third group connects and says nothing, often because a predictive dialer has no agent ready. Those systems leave different traces, yet a phone's “AI voice” warning may compress all of them into one confident label.

A new preprint from researchers at Scam AI shows why that shortcut is risky. Their interactive voice honeypot used language-model personas on real U.S. phone numbers and recorded 10,987 unwanted inbound calls over 66 days. Eleven days were excluded because the honeypot itself failed to greet callers. That left 7,233 normally greeted calls for the main breakdown.

Of those calls, 13.8% opened with a recording found on at least one other call. Another 13.1% opened with fresh audio that a commercial detector labeled synthetic. Together, those categories put machine-voiced openings at 26.9% in this honeypot. Another 9.9% stayed silent after the greeting, which the authors interpret as likely machine-placed, while 9.0% could not be scored.

The number is eye-catching. The detector audit is more useful.

One detector label covers two different mechanisms

The commercial detector labeled 29.3% of 6,192 scored openings as synthetic. Nearly half of those flags, 45%, belonged to replayed audio. The detector was often reacting to a file used on multiple calls, not necessarily to speech generated afresh during each conversation. That file could have been recorded from a person or synthesized once. The replay test cannot decide which.

The authors also found 2,756 pairs containing the same waveform. In 13.6% of those pairs, the two copies landed on opposite sides of the detector's threshold. Codec changes, telephone channels, segmentation, or the detector itself can apparently change a verdict even when the underlying audio is identical.

Eleven blinded listeners reviewed flagged clips and called 54.4% of the reviewed clips synthetic. Reweighting for the score distribution reduced that figure to 53.2%. Neither number is the detector's accuracy. The listening pool was drawn from audio the detector had already flagged and included no known-human negative class. It can estimate agreement or precision within that flagged group, but not recall, false-positive rate, or prevalence across the full corpus. Only 23 unflagged clips had listener judgments in this version.

The distinction between a large collection and a trusted label set also shaped Vaani's verified tier. Here, the researchers are unusually clear about which labels remain provisional. That candor makes the result more useful, not less.

Phone numbers vanish faster than campaigns

Most call blocking treats the originating number as the thing to track. The study found that a number is often the cheapest part of an operation to replace. One campaign made 68 calls from 68 numbers over 31 days. Across the corpus, a typical number within a campaign placed all its calls in one day, while the opening script continued for roughly two more weeks.

What persisted sat above the number. One recorded compliance notice appeared in six campaigns. One synthetic voice appeared in nine. By the time a number gathers enough complaints to acquire a bad reputation, the campaign may already have dropped it.

Scripts and behavior provide another view. Replayed calls talked over the honeypot more often and replayed a line within the same call more often than fresh synthetic or human-labeled calls. Machine-labeled calls also ignored more of the persona's questions. These are observations, not a finished classifier. They point toward a bundle of evidence: caller-number history, repeated opening text, waveform reuse, turn-taking, self-repetition, and a synthetic-speech score whose limits are recorded.

A detector describing its own output has a familiar cousin: a model explaining its own behavior. Direct model self-reports understated harmful behavior in a separate study. In both cases, the confident internal label is weak evidence without an external check.

The rules reach ordinary telemarketing too

In February 2024, the FCC ruled that AI-generated voices qualify as “artificial” voices under the Telephone Consumer Protection Act. The declaratory ruling applied existing consent, identification, and opt-out requirements. It did not ban every use of synthetic speech.

The FCC later proposed additional disclosure rules, including a notice at the beginning of a call using an AI-generated voice. In the honeypot corpus, 0.44% of calls whose opening was labeled synthetic disclosed automation. The paper treats that as a baseline, not a compliance rate. The disclosure rule was proposed, and the honeypot could see only part of the consent chain.

The calls were also less cinematic than the public debate around voice cloning. The detector labeled 33.8% of lead-generation spam synthetic, compared with 21.1% of calls the study classified as fraud. The FCC's 2024 action was publicly framed around impersonation and cloned voices, but synthetic speech in this sample appeared most often in routine telemarketing machinery.

Unwanted calling is already a large consumer problem. The FTC's 2025 Do Not Call Data Book records more than 2.6 million complaints, including 1,601,611 categorized as robocalls. The FTC does not independently verify each complaint, and those records do not identify synthetic voices. They establish the scale of the broader problem, not the prevalence claimed by this honeypot study.

A better response than guessing from the voice

If an unexpected caller asks for credentials, payment, or urgent action, the voice's naturalness should count for almost nothing. Hang up. Contact the organization through a number on its official app, website, bill, or payment card. Both the claimed identity and the voice arrived from the caller.

For call-screening software, the implementation choices are fairly concrete:

  1. Separate replay detection from synthetic-speech detection because they answer different questions.
  2. Store a detector's threshold, audio window, codec, and model version with its score. The score is evidence, not a verdict.
  3. Track repeated scripts and call behavior across changing numbers, with strict retention limits and privacy controls.
  4. Explain warnings with observable reasons, such as repeated wording or an unverified identity, instead of presenting “AI detected” as certainty.

An agent built with api.ish.chat can summarize those signals or guide a verification flow. It should not turn a black-box detector label into fact. ish.chat is useful here as a reasoning layer around provenance and user-controlled checks, not as a substitute for either.

The paper's limits are substantial. Its calls came from one fabricated lead identity actively seeded into a narrow set of lead-generation forms. The commercial detector is closed, the listener study is incomplete, and the analysis covers a call's opening rather than every later voice transition. The 26.9% figure does not describe all unwanted U.S. calls.

Its most durable result is about the unit of detection. Caller ID expires quickly; scripts, recordings, and voice assets do not. A screen that follows those persistent parts may still recognize a campaign after its latest number has disappeared.

#synthetic speech#robocalls#voice security#consumer protection#AI detection
Advertisement

Keep reading

Related stories

Browse the archive