用AI实时识骗反诈,保护隐私还高效。
AI-in-the-Loop: Privacy Preserving Real-Time Scam Detection and Conversational Scambaiting by Leveraging LLMs and Federated Learning
- 结合大模型与联邦学习,实时检测并打断诈骗对话。
- 生成对话流畅(困惑度22.3),人类评估效果优于基线。
- 支持隐私保护下的持续更新,适合安全敏感场景。
利用实时社交工程(如钓鱼、冒充、电话诈骗)的骗局仍是数字平台持续且不断演化的威胁。现有防御多为被动响应,难以在交互过程中提供有效防护。本文提出一种隐私保护的AI-in-the-loop框架,可主动实时检测并中断诈骗对话。系统融合指令微调的AI与安全感知的效用函数,在保持对话参与度的同时最小化伤害,并采用联邦学习实现模型持续更新而无需共享原始数据。实验表明,系统生成对话流畅(困惑度低至22.3,参与度≈0.80),人类评估显示其在真实感、安全性与有效性上显著优于强基线。在联邦设置下,使用FedAvg训练的模型可持续30轮更新,同时保持高参与度(≈0.80)、强相关性(≈0.74)及低个人身份信息泄露(≤0.0085)。即使引入差分隐私,新颖性与安全性仍稳定,表明高性能与强隐私可兼得。对守卫模型(LlamaGuard, LlamaGuard2/3, MD-Judge)的评估显示:更严格的审查降低信息暴露风险,但限制互动深度;更宽松设置则提升对话丰富性与检测能力,但增加隐私风险。据我们所知,这是首个将实时反诈、联邦隐私保护与可调安全调控统一于主动防御范式中的框架。
原文摘要 · Abstract (English)
Scams exploiting real-time social engineering -- such as phishing, impersonation, and phone fraud -- remain a persistent and evolving threat across digital platforms. Existing defenses are largely reactive, offering limited protection during active interactions. We propose a privacy-preserving, AI-in-the-loop framework that proactively detects and disrupts scam conversations in real time. The system combines instruction-tuned artificial intelligence with a safety-aware utility function that balances engagement with harm minimization, and employs federated learning to enable continual model updates without raw data sharing. Experimental evaluations show that the system produces fluent and engaging responses (perplexity as low as 22.3, engagement $\approx$0.80), while human studies confirm significant gains in realism, safety, and effectiveness over strong baselines. In federated settings, models trained with FedAvg sustain up to 30 rounds while preserving high engagement ($\approx$0.80), strong relevance ($\approx$0.74), and low PII leakage ($\leq$0.0085). Even with differential privacy, novelty and safety remain stable, indicating that robust privacy can be achieved without sacrificing performance. The evaluation of guard models (LlamaGuard, LlamaGuard2/3, MD-Judge) shows a straightforward pattern: stricter moderation settings reduce the chance of exposing personal information, but they also limit how much the model engages in conversation. In contrast, more relaxed settings allow longer and richer interactions, which improve scam detection, but at the cost of higher privacy risk. To our knowledge, this is the first framework to unify real-time scam-baiting, federated privacy preservation, and calibrated safety moderation into a proactive defense paradigm.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。