音频溯源模型在干净数据上表现好,但压缩后性能暴跌,不能只看干净准确率。
Clean Accuracy Does Not Guarantee Provenance Robustness: A Prospective Codec-Stress Evaluation of Audio Attribution
- 在单阶段编解码器传输后,用预设保真度元数据固定分析区域进行溯源测试。
- 两种模型在支持条件下损失高达70.3和61.0宏F1点,部分条件下降超50点。
- 编解码器影响显著,不同压缩率下感知指标与性能排序不一致,适合部署评估者看。
音频溯源——判断合成语音由哪个系统生成——在干净基准上已接近天花板准确率,但实际到达分析师的音频通常经过转码。本文报告了前瞻性注册的封闭集溯源测试,考察单阶段编解码器传输后的性能。分析区域基于编码前的保真度元数据固定,且在任何溯源模型训练前确定。在两个数据集上,WavLM-Base+ 的宏F1损失分别为53.5 [43.5, 63.6] 和70.3 [63.0, 77.5],而W2V2-BERT 2.0为61.0 [56.8, 65.1] 和49.8 [41.6, 57.9],所有结果均在同时性分量级带宽下测得。性能退化强烈依赖于条件与表示方式:同一支持网格内,WavLM损失范围从-0.4到+53.5点,两编码器在十二种条件中有六种超出预设±5点差异。即使使用经清洁数据验证的ECAPA-TDNN与代理锚点头,退化仍显著,表明问题不限于特定表示或弱线性头。该网格内匹配保真度对比不可估计;波形与感知指标对条件排序不同:MP3 8 kbit/s在SI-SDR中居中,在PESQ-WB中垫底,却导致最大损失。对于所测试的任务、数据集、表示和编解码器网格,仅凭干净准确率无法刻画部署鲁棒性。
原文摘要 · Abstract (English)
Audio provenance attribution - which system produced a synthetic utterance - is reported at near-ceiling accuracy on clean benchmarks, yet audio reaching an analyst has usually been transcoded. We report a prospectively registered measurement of closed-set attribution after single-stage codec transport, with the analysis region fixed from fidelity metadata before any attribution model was trained. On two corpora, in-support losses reach 53.5 [43.5, 63.6] and 70.3 [63.0, 77.5] Macro-F1 points for WavLM-Base+, and 61.0 [56.8, 65.1] and 49.8 [41.6, 57.9] for W2V2-BERT 2.0, under simultaneous component-level bands. Degradation is strongly condition- and representation-dependent: within one in-support grid WavLM losses run from -0.4 to +53.5 points, and the two encoders differ beyond a prespecified +/-5-point margin at six of twelve conditions. A clean-qualified ECAPA-TDNN and a Proxy-Anchor head degrade comparably, so the effect is not confined to one representation family or a weak linear head. The registered matched-fidelity comparison was not estimable on this grid, and waveform and perceptual measures order the conditions differently: MP3 at 8 kbit/s ranks mid-grid on SI-SDR but last on PESQ-WB while causing the largest loss. For the tested tasks, corpora, representations and codec grid, a clean accuracy figure does not by itself characterise deployment robustness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。