arXiv:2605.17737cs.SD2026-05中稿 · IJCAI被引 2

通过声学指纹识别说话人独特发音模式,有效检测语音伪造

Profiling the Voice: Speaker-Specific Phoneme Fingerprinting for Speech Deepfake Detection

论文配图:Profiling the Voice: Speaker-Specific Phoneme Fingerprinting for Speech Deepfake Detection
图 1 · 摘自论文原文
  • 用轻量GMM建模说话人特有音素发音习惯,无需伪造数据训练
  • 在中文名人语音伪造检测中误报率降低37.2%,优于现有方法
  • 提供音素级可解释性,适合司法鉴定与高危人物防护场景

生成式AI的快速发展使语音伪造越来越难以与真实人声区分,对公众人物等重要目标构成严重威胁。现有检测系统多依赖通用黑箱模型,无法捕捉说话人特有的发音特征,且缺乏可解释性。本文提出基于音素的语音画像(PVP)框架,将检测范式从宏观语句分析转向微观音素建模,通过仅使用真实参考语音估计轻量级高斯混合模型(GMM),捕获特定说话人惯常的发音声学分布。该设计实现高效建模,并能泛化至未见过的伪造攻击,无需大量伪造样本训练。此外,我们构建了首个大规模中文公众人物语音伪造数据集用于基准测试。实验表明,PVP在公众人物伪造场景下显著优于现有通用检测器,误报率(EER)大幅降低,同时提供音素级别的可解释性,助力司法取证。代码与数据已公开于:https://github.com/JunXue-tech/PVP

原文摘要 · Abstract (English)

The rapid advancement of generative AI has made audio deepfakes increasingly indistinguishable from authentic human vocals, posing significant threats to persons-of-interest (POI) such as public figures. Current detection systems primarily rely on generic, black-box models that fail to capture speaker-specific idiosyncratic traits and lack interpretability. In this paper, we propose Phoneme-based Voice Profiling (PVP), a novel personalized defense framework. By shifting the detection paradigm from macro-utterance analysis to micro-phonetic modeling, PVP captures the unique acoustic distributions underlying a POI's habitual articulatory patterns. Specifically, our framework models speaker-specific phonetic realizations using lightweight Gaussian Mixture Models (GMMs) estimated solely from bona fide reference speech. This design enables data-efficient profiling and robust generalization to previously unseen spoofing attacks without requiring heavy spoof-specific training. Furthermore, we introduce the first large-scale Chinese POI deepfake dataset to benchmark speaker-specific detection. Experimental results demonstrate that PVP significantly outperforms state-of-the-art generic detectors in POI spoofing scenarios, achieving substantial EER reductions while providing fine-grained, phoneme-level interpretability for forensic analysis. Code and data are available at: https://github.com/JunXue-tech/PVP

语音伪造声纹识别可解释性中文数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。