用类人推理提升语音伪造检测的泛化与可解释性
Towards Robust Speech Deepfake Detection via Human-Inspired Reasoning
- 结合大音频语言模型与人类标注的思维链数据
- 在多个新域上检测准确率显著优于现有方法
- 能给出符合人类感知的判断理由,适合安全审核场景
当前生成式音频模型可能被用于非法冒充他人获取隐私信息。为应对这一问题,语音深伪检测(SDD)方法应运而生。然而,现有SDD方法普遍存在对新音频领域和生成器泛化能力差、缺乏可解释性的问题,尤其缺少类人推理能力来自然解释判断依据并提供人类可感知的线索。本文提出HIR-SDD框架,融合大音频语言模型(LALMs)与基于全新人工标注数据集构建的思维链推理机制。实验表明,该方法不仅有效提升检测性能,还能为预测结果提供合理且可理解的解释。
原文摘要 · Abstract (English)
The modern generative audio models can be used by an adversary in an unlawful manner, specifically, to impersonate other people to gain access to private information. To mitigate this issue, speech deepfake detection (SDD) methods started to evolve. Unfortunately, current SDD methods generally suffer from the lack of generalization to new audio domains and generators. More than that, they lack interpretability, especially human-like reasoning that would naturally explain the attribution of a given audio to the bona fide or spoof class and provide human-perceptible cues. In this paper, we propose HIR-SDD, a novel SDD framework that combines the strengths of Large Audio Language Models (LALMs) with the chain-of-thought reasoning derived from the novel proposed human-annotated dataset. Experimental evaluation demonstrates both the effectiveness of the proposed method and its ability to provide reasonable justifications for predictions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。