arXiv:2504.05831cs.CL2025-04NeurIPS被引 2

用鲁棒优化提升大模型在分布偏移下的对齐能力

Leveraging Robust Optimization for LLM Alignment under Distribution Shifts

  • 通过分类器为样本赋校准值,识别与人类偏好分布的匹配度
  • 在关键数据区域最小化最坏情况损失,提升对齐效果
  • 适合关注大模型价值观对齐与分布鲁棒性的研究者

偏好对齐方法对引导大语言模型生成符合人类价值观的输出至关重要。尽管近期方法常依赖由大模型生成的合成数据以实现可扩展性和成本效益,但这种依赖可能引入分布偏移,削弱对人类偏好细微特征的表征能力。本文提出一种新的分布感知优化框架,在存在分布偏移时仍能改进偏好对齐。该方法首先利用训练良好的分类器为每个训练样本分配校准值,量化其与目标人类偏好分布的匹配程度。这些值随后被纳入鲁棒优化目标,最小化数据空间中与人类偏好最相关的区域的最坏情况损失。通过明确聚焦于目标分布,该方法缓解了分布不匹配的影响,提升了生成更符合预期价值回应的能力。

原文摘要 · Abstract (English)

Preference alignment methods are increasingly critical for steering large language models (LLMs) to generate outputs consistent with human values. While recent approaches often rely on synthetic data generated by LLMs for scalability and cost-efficiency reasons, this reliance can introduce distribution shifts that undermine the nuanced representation of human preferences needed for desirable outputs. In this paper, we propose a novel distribution-aware optimization framework that improves preference alignment despite such shifts. Our approach first leverages well-learned classifiers to assign a calibration value to each training sample, quantifying its alignment with the target human-preferred distribution. These values are then incorporated into a robust optimization objective that minimizes the worst-case loss over regions of the data space most relevant to human preferences. By explicitly focusing optimization on the target distribution, our approach mitigates the impact of distributional mismatch and improves the generation of responses that better reflect intended values.

大模型对齐鲁棒优化分布偏移

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。