对比不同训练目标对语音伪造检测适配器几何结构的影响。
How Meta-Learning Shapes LoRA Adapter Geometry in Speech Deepfake Detection

- 固定模型与适配器结构,仅改变训练目标,分析其对适配器几何的影响。
- MLDG使关键投影层的更新更集中于查询与键,输出层更分散。
- 该几何差异解释了域泛化性能差距,适用于关注适配器设计的研究者。
在相同的冻结自监督语音模型上训练低秩适配器时,元学习用于领域泛化(MLDG)相比经验风险最小化(ERM)能提升分布外语音伪造检测性能。由于架构和适配器容量固定,性能差距源于训练目标对适配器形状的塑造方式不同。本文提出一种诊断方法:在保持架构、秩、数据和随机种子一致的前提下,仅改变训练目标,通过完成后的适配器上的经验费舍尔矩阵来比较ERM与MLDG所形成的几何结构。采用有效秩诊断,按投影和深度分解适配器变化中哪些影响损失。结果表明,两种目标并非同等重塑所有投影:损失相关更新集中在查询与键投影,而输出投影的更新更分散,这一现象在六个语料库中一致出现,且在高层最显著。该差异在合并更新中同样存在,不受低秩分解影响,说明其反映的是有效更新的几何本质。研究揭示,ERM与MLDG的性能差异不仅体现在误差率,更在于损失相关容量在适配器内的组织方式,而损失感知的适配器几何可作为理解此差异的新视角。
原文摘要 · Abstract (English)
Meta-learning for domain generalization (MLDG) improves out-of-distribution speech deepfake detection over empirical risk minimization (ERM) when both objectives train low-rank adapters on the same frozen self-supervised speech model. Because the architecture and adapter capacity are held fixed, this gap points to differences in how the training objective shapes the adapter, yet the field characterizes objectives through error rates rather than through the geometry of the solution they reach. We introduce a descriptive diagnostic for this question: holding architecture, rank, data, and seeds fixed and varying only the objective, we use the empirical Fisher on the finished adapter to compare the geometry that ERM and MLDG leave behind. We characterize each adapter with effective-rank diagnostics that separate where the adapter changes from where those changes matter to the loss, resolved by projection and by depth. Applied to ERM and MLDG, the diagnostic shows that the objective does not reshape all adapter projections alike: the loss-relevant update concentrates in the query and key projections while becoming more distributed in the output projection, consistently across six corpora and most strongly in the upper layers. The same contrast appears in the merged update independently of the low-rank factorization, indicating that it reflects the geometry of the effective update rather than the parameterization. These results show that the gap between ERM and MLDG is not only a difference in error rate, but a difference in how loss-relevant capacity is organized inside the adapter, and that loss-aware adapter geometry is a way to see it.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。