用动态几何锚点提升大模型对齐的鲁棒性
Learning Where It Matters: Geometric Anchoring for Robust Preference Alignment
- 用当前策略的局部对抗扰动作为动态参考基准
- 通过锚点差距自适应降低敏感样本权重
- 适合在噪声数据下训练稳定可靠的对齐模型
直接偏好优化(DPO)等方法通过固定参考策略对齐大语言模型,但随着策略漂移,静态参考会失准,导致分布不匹配并放大噪声监督中的虚假信号。而无参考方法则易出现奖励漂移。本文提出几何锚点偏好优化(GAPO),以当前策略的小邻域内对抗扰动作为动态、几何感知的锚点,形成悲观基线。该锚点支持自适应重加权机制,根据每对偏好在局部的敏感度调节其重要性。进一步引入锚点差距(Anchor Gap),即策略与锚点间的奖励差异,在平滑条件下可近似最坏情况下的局部边界退化。优化基于该差距加权的逻辑回归目标,可削弱几何脆弱实例的影响,强化稳健偏好信号。在多种噪声设置下,GAPO均显著提升鲁棒性,且在标准大模型对齐与推理基准上表现持平或更优。
原文摘要 · Abstract (English)
Direct Preference Optimization (DPO) and related methods align large language models from pairwise preferences by regularizing updates against a fixed reference policy. As the policy drifts, a static reference, however, can become increasingly miscalibrated, leading to distributional mismatch and amplifying spurious preference signals under noisy supervision. Conversely, reference-free variants avoid mismatch but often suffer from unconstrained reward drift. We propose Geometric Anchor Preference Optimization (GAPO), which replaces the fixed reference with a dynamic, geometry-aware anchor: an adversarial local perturbation of the current policy within a small radius that serves as a pessimistic baseline. This anchor enables an adaptive reweighting mechanism, modulating the importance of each preference pair based on its local sensitivity. We further introduce the Anchor Gap, the reward discrepancy between the policy and its anchor, and show under smoothness conditions that it approximates worst-case local margin degradation. Optimizing a logistic objective weighted by this gap downweights geometrically brittle instances while emphasizing robust preference signals. Across diverse noise settings, GAPO consistently improves robustness while matching or improving performance on standard LLM alignment and reasoning benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。