提出隐私与抗干扰并重的语言模型对齐理论新边界,突破传统认知。
Improved Bounds for Private and Robust Alignment
- 基于对数损失与最大似然思想,实现近优隐私对齐
- 首次建立联合隐私与对抗攻击下的在线对齐理论界限
- 发现现有离线算法实际性能优于已知上限,适用于高鲁棒性场景
本文从理论角度研究语言模型在隐私约束和对抗污染双重条件下的对齐问题,推导出离线与在线设置下次优性差距的上界。考虑偏好标签同时受隐私保护限制和对抗性污染影响,分析了隐私优先与污染优先两种不同机制。在仅隐私约束情形下,证明对数损失结合极大似然风格算法可达到近最优率,颠覆传统认知;在联合隐私与污染情形下,发现现有离线算法实际上同时提供更强的抗污染与隐私保障,进而改进了纯污染场景下的理论界限。此外,首次给出了私密且鲁棒在线对齐的理论结果。这些成果依赖于在隐私与污染条件下对对数损失和平方损失的新统一收敛性保证,我们认为其在学习理论与统计学中具有广泛适用性。
原文摘要 · Abstract (English)
In this paper, we study the private and robust alignment of language models from a theoretical perspective by establishing upper bounds on the suboptimality gap in both offline and online settings. We consider preference labels subject to privacy constraints and/or adversarial corruption, and analyze two distinct interplays between them: privacy-first and corruption-first. For the privacy-only setting, we show that log loss with an MLE-style algorithm achieves near-optimal rates, in contrast to conventional wisdom. For the joint privacy-and-corruption setting, we first demonstrate that existing offline algorithms in fact provide stronger guarantees -- simultaneously in terms of corruption level and privacy parameters -- than previously known, which further yields improved bounds in the corruption-only regime. In addition, we also present the first set of results for private and robust online alignment. Our results are enabled by new uniform convergence guarantees for log loss and square loss under privacy and corruption, which we believe have broad applicability across learning theory and statistics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。