arXiv:2602.18868math.OCcs.LG2026-02被引 5

提出新方法控制模型微调速度,揭示其理论极限。

Limits of Convergence-Rate Control for Open-Weight Safety

  • 通过权重谱结构重参数化,实现优化速度的可证明控制
  • SpecDef算法在非对抗场景下有效减缓一阶与二阶优化
  • 发现攻击者可线性扩大模型规模恢复快速收敛,突破现有控制上限

开放权重的基础模型在发布后可能被微调用于有害目的,而现有训练抗性方法缺乏理论保障。将此类干预视为收敛速率控制问题,可将优化速度与模型权重的谱结构关联。基于此洞察,我们提出一种通过谱重参数化实现收敛速率控制的新理解,并推导出名为SpecDef的算法,在非对抗设置下可同时理论和实证地减缓一阶与二阶优化。在对抗设置中,我们建立了广泛类别的收敛速率控制方法(包括本文方法)的根本极限:拥有足够知识的攻击者可通过线性增加模型规模恢复快速收敛。因此,未来工作需探索不等价于收敛速率控制的方法。

原文摘要 · Abstract (English)

Open-weight foundation models can be fine-tuned for harmful purposes after release, yet no existing training resistance methods provide theoretical guarantees. Treating these interventions as convergence-rate control problems allows us to connect optimization speed to the spectral structure of model weights. We leverage this insight to develop a novel understanding of convergence rate control through spectral reparameterization and derive an algorithm, SpecDef, that can both provably and empirically slow first- and second-order optimization in non-adversarial settings. In adversarial settings, we establish a fundamental limit on a broad class of convergence rate control methods including our own: an attacker with sufficient knowledge can restore fast convergence at a linear increase in model size. In order to overcome this limitation, future works will need to investigate methods that are not equivalent to controlling convergence rate.

模型安全收敛控制谱分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。