用语音发音片段特征检测深度伪造音频,更难被模型模仿。
Forensic deepfake audio detection using segmental speech features
- 利用发音部位相关的语音片段特征进行检测
- 部分发音特征可有效识别深度伪造音频
- 适合需要逐案分析的司法鉴定场景
本研究探索使用语音片段声学特征检测深度伪造音频的潜力。这些特征因与人类发音生理过程密切相关,具有高度可解释性,且更难被深度伪造模型复现。实验表明,某些常用于司法语音比对(FVC)的片段特征能有效识别深度伪造音频,而部分全局特征则作用有限。结果强调需采用有别于传统FVC的方法进行音频深度伪造检测,并提出一种针对特定说话人的检测框架,该框架不同于当前主流的说话人无关系统。在司法鉴定中,逐案可解释性和对个体发音差异的敏感性至关重要,而说话人特定方法正具备这一优势。
原文摘要 · Abstract (English)
This study explores the potential of using acoustic features of segmental speech sounds to detect deepfake audio. These features are highly interpretable because of their close relationship with human articulatory processes and are expected to be more difficult for deepfake models to replicate. The results demonstrate that certain segmental features commonly used in forensic voice comparison (FVC) are effective in identifying deep-fakes, whereas some global features provide little value. These findings underscore the need to approach audio deepfake detection using methods that are distinct from those employed in traditional FVC, and offer a new perspective on leveraging segmental features for this purpose. In addition, the present study proposes a speaker-specific framework for deepfake detection, which differs fundamentally from the speaker-independent systems that dominate current benchmarks. While speaker-independent frameworks aim at broad generalization, the speaker-specific approach offers advantages in forensic contexts where case-by-case interpretability and sensitivity to individual phonetic realization are essential.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。