通过分析情绪与声学特征的跨层次不一致,提升音频伪造检测能力
Emotion and Acoustics Should Agree: Cross-Level Inconsistency Analysis for Audio Deepfake Detection
- 将情绪与声学特征映射到统一空间,捕捉多粒度不一致
- 在ASVspoof 2019LA和2021LA上实现更优检测性能
- 适合关注音频伪造检测中细粒度异常的研学者
音频伪造检测(ADD)旨在区分真实语音与伪造语音。以往研究通常假设声学与情绪特征间强相关即为真实,聚焦于增强或测量此类相关性。然而现有方法常孤立处理两类特征,依赖相关性指标,忽略了二者间细微的时序错位及突变断点。为此,本文提出EAI-ADD,将跨层次情绪-声学不一致作为核心检测信号。首先将情绪与声学表征投影至可比空间;随后逐级融合帧级与语句级情绪特征与声学特征,以捕捉不同时间粒度下的不一致性。在ASVspoof 2019LA和2021LA数据集上的实验表明,所提方法优于基线,为音频反伪造检测提供了更有效的解决方案。
原文摘要 · Abstract (English)
Audio Deepfake Detection (ADD) aims to detect spoof speech from bonafide speech. Most prior studies assume that stronger correlations within or across acoustic and emotional features imply authenticity, and thus focus on enhancing or measuring such correlations. However, existing methods often treat acoustic and emotional features in isolation or rely on correlation metrics, which overlook subtle desynchronization between them and smooth out abrupt discontinuities. To address these issues, we propose EAI-ADD, which treats cross level emotion acoustic inconsistency as the primary detection signal. We first project emotional and acoustic representations into a comparable space. Then we progressively integrate frame level and utterance level emotion features with acoustic features to capture cross level emotion acoustic inconsistencies across different temporal granularities. Experimental results on the ASVspoof 2019LA and 2021LA datasets demonstrate that the proposed EAI-ADD outperforms baselines, providing a more effective solution for audio anti spoofing detection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。