通过分层决策融合提升假音频检测准确率,跨数据集表现最优。
Layer-Wise Decision Fusion for Fake Audio Detection Using XLS-R

- 先分层决策再融合,避免特征坍塌
- 在In-the-Wild数据集上误报率仅6.90%
- 模型可解释性强,便于分析决策机制
近期假音频检测方法常利用大型语音模型获取鲁棒的语音表征。这些模型通常深度很深,提供多层表征。然而,现有方法多依赖单一层次表征或特征融合,生成全局表征进行决策,容易忽略多层丰富信息,可能导致特征坍塌。本文提出一种新型分层决策融合方法,在每层独立决策后进行融合,相比其他强基线,在In-the-Wild数据集上达到最低的等错误率(EER 6.90%),实现最佳跨数据集性能。该模型设计还增强了可解释性,使我们能够深入分析决策背后的机制。
原文摘要 · Abstract (English)
Recent fake audio detection methods often leverage large speech models to achieve robust speech representations. These models are typically very deep, providing multiple layer-wise representations. However, current works often rely solely on single layer representation or feature fusion to extract one utterance-level representation for decision making. These methods risk underutilizing rich information from multiple layers and might induce feature collapse. We propose a novel layer-wise decision fusion method that applies fusion after per-layer decision making and achieves the best cross-dataset performance on In-the-Wild dataset (EER 6.90%) compared to other strong baselines. Our model design also makes the model more transparent, allowing us to conduct detailed analysis to reveal the underlying mechanism of decision making.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。