分析1万通真实诈骗电话,揭示骗子话术套路与行为规律。
Anatomy of a Scam Call: What 10,000 real scam and spam calls reveal about how phone scammers operate

- 用AI语音陷阱收集超1万通诈骗来电,还原真实通话流程。
- 骗子更关注地址和生日,而非支付信息;年龄越大越耗时,但要的东西不变。
- 开场白就能预测是否升级诈骗,简单模型已可有效识别。
电话诈骗普遍且代价高昂,但其运作机制极少被大规模观察。本研究分析了由AI语音代理蜜罐在54天内收集的完整语料库,共10,211通入站诈骗与垃圾电话,涵盖913小时音频和330,956次转录对话,来自5,780个不同号码。区分了直接诈骗(索取敏感信息)与合法但具有掠夺性的线索生成(“垃圾”电话)。诈骗活动有明确办公时间(工作日通话量是周末的6.6倍);数以千计的临时号码重复使用有限脚本(三十种开场模板,前五种占一半流量);且更常索要身份锚点(如住址、出生日期),通过持续施压和伪装权威而非威胁。核心实验表明:目标年龄每增加十年,骗子平均多消耗15%对话轮次(率比1.15,95%置信区间1.08-1.23,随机化p=0.005),但请求内容不变(26.3%的通话进入敏感信息索取阶段,每十年几率比0.99,95%置信区间0.90-1.08)。第二项实验将早期检测作为基准:仅从首次发言即可预测升级概率,第一句时ROC-AUC达0.72,第八句提升至0.87;朴素词袋模型表现媲美微调的本地语言模型。电话诈骗呈现出高度模板化的产业特征——针对目标投入力度随年龄变化,但目标始终不变。
原文摘要 · Abstract (English)
Telephone fraud is pervasive and costly, but its inner workings are rarely observed at scale. We analyze a complete corpus of 10,211 inbound scam and spam calls -- 913 hours of audio and 330,956 transcribed turns from 5,780 distinct numbers -- collected over 54 days by an AI voice-agent honeypot that answered callers and kept them talking, and introduced in a companion data descriptor. We separate outright scams, which solicit sensitive information, from the larger stream of predatory but legal lead generation ("spam") that feeds them. Scam operations keep office hours (6.6x more calls per weekday than weekend day); thousands of disposable numbers run a small catalog of recycled scripts (thirty opening clusters, half the traffic in the top five); and callers solicit identity anchors -- a home address and a date of birth -- far more often than payment credentials, pressing through persistence and manufactured authority rather than overt threats. Our central experiment asks: does it matter who picks up? Every seeded lead carried one of ten fictitious identities drawn uniformly at random, so the identity a fraud operation reaches is fixed before the caller exists. Across 1,823 randomized calls, scammers spent about 15% more conversational turns per decade of the target's apparent age (rate ratio 1.15, 95% CI 1.08-1.23; randomization p = 0.005) -- yet what they asked for did not change (26.3% of calls reached a request for sensitive information; odds ratio 0.99 per decade, 95% CI 0.90-1.08). A second experiment casts early detection as a benchmark: from a scammer's opening lines alone, on a caller-disjoint split, escalation is predictable at 0.72 ROC-AUC from the first line and 0.87 by the eighth, and a plain bag-of-words classifier matches a fine-tuned on-device language model. Telephone fraud emerges as a templated industry that varies how hard it works a target, but not what it wants.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。