arXiv:2609.07079cs.SDcs.CL2026-09

对比两种语音诈骗检测方法,发现无监督模型更适配真实场景

Comparing Self-Supervised and Domain-Invariant Features for Cross-Domain Voice Phishing Detection

  • 用跨域评估比较声调特征与自监督模型性能
  • 自监督模型最高达94.2%准确率,但需真实样本
  • 声调特征支持零样本部署,适合隐私受限场景

语音诈骗检测面临三大挑战:真实犯罪录音因隐私限制不可得;即便获得,样本量极少,不足以微调;且需要轻量级纯声学检测方案,作为大型自监督模型的替代。本文通过跨域评估,比较了领域不变的韵律特征与自监督表示(HuBERT、wav2vec2.0),训练使用基于场景的演员录音,测试使用真实犯罪通话。领域不变的韵律特征实现69.5% F1(零样本)和71.0% F1(5样本学习)。HuBERT在5样本下达到最高性能(94.2% F1),而wav2vec2.0表现出高精度导向特性(90.2% F1,99.4% 精确率,5样本)。研究揭示根本权衡:领域不变特征可在无真实数据时实现零样本部署,而自监督方法性能更高,但需真实样本和计算资源。

原文摘要 · Abstract (English)

Voice phishing detection faces three critical challenges: real criminal recordings are unavailable due to privacy constraints; when available, only a handful of samples exist, insufficient for fine-tuning; and lightweight acoustic-only detection is needed as an alternative to large self-supervised models. We compare domain-invariant prosodic features and self-supervised representations (HuBERT, wav2vec2.0) through cross-domain evaluation-training on scenario-based actor recordings and testing on authentic criminal calls. Domain-invariant prosodic features achieve 69.5% F1 zero-shot and 71.0% with 5-shot learning. HuBERT achieves highest performance (94.2% F1, 5-shot), while wav2vec2.0 exhibits a precision-oriented detection profile (90.2% F1 with 99.4% precision, 5-shot). These findings reveal fundamental trade-offs: domain-invariant features enable zero-shot deployment when no real data exists, while SSL methods achieve higher performance but require real samples and compute.

语音诈骗自监督零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。