通过参数回归解析音频伪造源,实现对未知编码器的精细识别。
Towards Neural Audio Codec Source Parsing
- 将音频伪造源识别转为生成参数的结构化回归任务
- 在多个基准数据集上超越欧氏空间基线模型性能
- 适用于需要细粒度溯源的数字取证场景
一种新型音频深度伪造——编码器伪造(CFs)近期受到关注,由利用神经音频编码器(NACs)的语音语言模型生成。为应对这一挑战,社区已建立专用基准并提出针对性检测策略。随着研究深入,检测目标已从二分类转向源归属,包括开集归属,旨在识别生成所用的NAC并标记推理中出现的未知新编码器。该转变提升了司法溯源的可解释性与责任追溯能力。然而,开集归属仍存在根本局限:虽可检测出不熟悉编码器,却无法刻画或识别具体未见编码器。它将此类输入视为通用“未知”,缺乏对其内部配置的洞察。这导致严重缺陷:对新NAC泛化能力有限,且无法分辨同一编码器家族内的细微差异。为此,我们提出神经音频编码器源解析(NACSP)——一种范式转变,将CFs的源归属重构为对量化器、带宽、采样率等生成式NAC参数的结构化回归。我们将NACSP建模为多任务回归任务以预测这些参数,并建立了首个基于多种先进语音预训练模型(PTMs)的综合性基准。为此,我们提出HYDRA框架,利用双曲几何从PTM表示中解耦复杂潜在属性。通过在多个曲率感知的双曲子空间上使用任务特定注意力,HYDRA实现了卓越的多任务泛化能力。大量实验表明,相较于运行于欧氏空间的基线方法,HYDRA在基准CFs数据集上取得最优结果。
原文摘要 · Abstract (English)
A new class of audio deepfakes-codecfakes (CFs)-has recently caught attention, synthesized by Audio Language Models that leverage neural audio codecs (NACs) in the backend. In response, the community has introduced dedicated benchmarks and tailored detection strategies. As the field advances, efforts have moved beyond binary detection toward source attribution, including open-set attribution, which aims to identify the NAC responsible for generation and flag novel, unseen ones during inference. This shift toward source attribution improves forensic interpretability and accountability. However, open-set attribution remains fundamentally limited: while it can detect that a NAC is unfamiliar, it cannot characterize or identify individual unseen codecs. It treats such inputs as generic ``unknowns'', lacking insight into their internal configuration. This leads to major shortcomings: limited generalization to new NACs and inability to resolve fine-grained variations within NAC families. To address these gaps, we propose Neural Audio Codec Source Parsing (NACSP) - a paradigm shift that reframes source attribution for CFs as structured regression over generative NAC parameters such as quantizers, bandwidth, and sampling rate. We formulate NACSP as a multi-task regression task for predicting these NAC parameters and establish the first comprehensive benchmark using various state-of-the-art speech pre-trained models (PTMs). To this end, we propose HYDRA, a novel framework that leverages hyperbolic geometry to disentangle complex latent properties from PTM representations. By employing task-specific attention over multiple curvature-aware hyperbolic subspaces, HYDRA enables superior multi-task generalization. Our extensive experiments show HYDRA achieves top results on benchmark CFs datasets compared to baselines operating in Euclidean space.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。