针对老年语音伪造检测难题,提出新框架与数据集,显著提升识别准确率。
Bridging the Age Gap: Towards Detecting Neural Audio Codec Synthesized Elderly Speech Deepfake
- 融合多模态基础模型,利用跨模态预训练增强对老年语音的感知能力
- 在中英文数据集上平均误报率仅1.66%,优于现有最先进方法
- 适合语音安全、数字身份验证等领域的研究者和开发者参考
本研究提出老年语音伪造检测(ECFD)任务,并发布中英文双语的老年人工音频伪造数据集(ECF)。实验表明,以往基于基准数据集训练的语音伪造检测器在老年语音上泛化性能差,暴露了关键漏洞。我们进一步假设并验证:如LanguageBind(LB)和ImageBind(IB)这类多模态基础模型,因其在跨模态预训练中接触过老年内容,对老年语音伪造更具敏感性。受此启发,我们探索多模态模型融合以提升性能,提出BONSAI框架,采用Jensen-Shannon散度作为融合机制。BONSAI结合LB与IB,在测试中实现平均错误率(EER)1.66%,超越单个模型及现有最优基线,为ECFD任务建立新基准。
原文摘要 · Abstract (English)
In this study, we introduce the Elderly CodecFake Detection (ECFD) task and release the Elderly-CodecFake (ECF) dataset in English and Chinese. We show that state-of-the-art CF detectors trained on previous benchmark CF datasets generalize poorly to elderly speech, revealing a critical vulnerability. We further hypothesize and demonstrate that multimodal foundation models (FMs) such as LanguageBind (LB) and ImageBind (IB) are more effective for ECFD due to their exposure to elderly content during cross-modal pretraining. Motivated by prior evidence that fusion of FMs enhances downstream performance, we explore fusion of FMs for ECFD. To this end, we propose BONSAI, a novel framework that employs Jensen-Shannon Divergence as the fusion mechanism. BONSAI with the fusion of LB and IB achieves an average EER (%) of 1.66 and outperforms individual FMs as well as competitive SOTA baselines, establishing a new benchmark for the ECFD task.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。