arXiv:2607.20023cs.SD2026-07中稿 · Interspeech 2025

通过分层决策融合提升假音频检测准确率,跨数据集表现最优。

Layer-Wise Decision Fusion for Fake Audio Detection Using XLS-R

论文配图:Layer-Wise Decision Fusion for Fake Audio Detection Using XLS-R
图 1 · 摘自论文原文
  • 先分层决策再融合,避免特征坍塌
  • 在In-the-Wild数据集上误报率仅6.90%
  • 模型可解释性强,便于分析决策机制

近期假音频检测方法常利用大型语音模型获取鲁棒的语音表征。这些模型通常深度很深,提供多层表征。然而,现有方法多依赖单一层次表征或特征融合,生成全局表征进行决策,容易忽略多层丰富信息,可能导致特征坍塌。本文提出一种新型分层决策融合方法,在每层独立决策后进行融合,相比其他强基线,在In-the-Wild数据集上达到最低的等错误率(EER 6.90%),实现最佳跨数据集性能。该模型设计还增强了可解释性,使我们能够深入分析决策背后的机制。

原文摘要 · Abstract (English)

Recent fake audio detection methods often leverage large speech models to achieve robust speech representations. These models are typically very deep, providing multiple layer-wise representations. However, current works often rely solely on single layer representation or feature fusion to extract one utterance-level representation for decision making. These methods risk underutilizing rich information from multiple layers and might induce feature collapse. We propose a novel layer-wise decision fusion method that applies fusion after per-layer decision making and achieves the best cross-dataset performance on In-the-Wild dataset (EER 6.90%) compared to other strong baselines. Our model design also makes the model more transparent, allowing us to conduct detailed analysis to reveal the underlying mechanism of decision making.

假音频检测分层融合XLS-R

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。