arXiv:2501.06514cs.SDcs.AI2025-01被引 16

提出开放集神经编解码溯源任务,解决音频深度伪造的来源追踪难题。

Neural Codec Source Tracing: Toward Comprehensive Attribution in Open-Set Condition

  • 定义开放集神经编解码溯源任务,支持未知源识别。
  • 构建包含11种生成方法的ST-Codecfake数据集,覆盖跨语言和异常样本。
  • 首次在开放集条件下评估溯源模型,揭示其对真实音频泛化能力不足。

当前音频深度伪造检测正从二分类转向多类溯源任务。然而现有研究仅限封闭集场景,未考虑开放集挑战。本文提出神经编解码溯源(NCST)任务,可实现开放集下的神经编解码器分类与可解释的ALM检测。我们构建了ST-Codecfake数据集,包含由11种先进神经编解码器生成的双语语音样本及基于ALM的分布外(OOD)测试样本。同时建立全面的溯源基准以评估模型在开放集条件下的表现。实验表明,尽管模型在分布内(ID)分类和OOD检测上表现良好,但在识别未见过的真实音频时仍缺乏鲁棒性。相关数据集与代码已公开。

原文摘要 · Abstract (English)

Current research in audio deepfake detection is gradually transitioning from binary classification to multi-class tasks, referred as audio deepfake source tracing task. However, existing studies on source tracing consider only closed-set scenarios and have not considered the challenges posed by open-set conditions. In this paper, we define the Neural Codec Source Tracing (NCST) task, which is capable of performing open-set neural codec classification and interpretable ALM detection. Specifically, we constructed the ST-Codecfake dataset for the NCST task, which includes bilingual audio samples generated by 11 state-of-the-art neural codec methods and ALM-based out-ofdistribution (OOD) test samples. Furthermore, we establish a comprehensive source tracing benchmark to assess NCST models in open-set conditions. The experimental results reveal that although the NCST models perform well in in-distribution (ID) classification and OOD detection, they lack robustness in classifying unseen real audio. The ST-codecfake dataset and code are available.

音频伪造溯源开放集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。