对比两种语音溯源方法,发现成对验证会降低精度。
The Hidden Cost of Pairwise Verification in Synthetic Speech Source Tracing

- 用全局锚定替代成对验证,提升溯源准确率
- 成对验证使嵌入方向变少,降低相似生成器区分度
- 适合关注语音溯源模型设计的研究者
开放集语音源溯源常被建模为验证问题,促使采用生物识别中的成对度量学习目标。我们在相同骨干网络和固定数据与训练轮次预算下,比较了全局锚定与成对验证在 MLAAD(域内)和 STOPA(域外)上的表现。实验显示,全局锚定的域内错误率(EER)为8.61%,显著低于成对方法(12-15% EER),即使使用对抗性挖掘和 XLS-R 微调。因成对目标直接优化相似性,导致方差集中在更少的嵌入方向上,削弱了相近生成器间的区分能力。为检验此效应,我们人为施加类似瓶颈至全局监督基线,其性能仍保持竞争力。结合嵌入空间分析($k_{99}$),结果表明误差差距并非仅由维度决定,而是成对目标对保留方向的塑造所致。
原文摘要 · Abstract (English)
Open-set source tracing is increasingly framed as a verification problem, motivating the use of pairwise metric-learning objectives from biometrics. We thus compare global anchoring and pairwise verification under matched backbones and a fixed data and epoch budget on MLAAD (in-domain) and STOPA (out-of-domain). In our runs, global anchoring yields lower in-domain error (8.61% EER) than pairwise variants (12-15% EER), even with rival mining and XLS-R finetuning. Because pairwise objectives optimize similarity directly, they concentrate variance into fewer embedding directions, reducing resolution among closely related generators. To test if this drives the drop, we impose a similar bottleneck to the globally supervised baseline, yet the baseline remains competitive. Together with an embedding-space analysis ($k_{99}$), these results suggest that the gap is not explained by dimensionality alone, but rather by the pairwise objective's shaping of the retained directions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。