arXiv:2509.12936cs.LGcs.CL2025-09Conference of the …被引 1

对比五种对齐方法,揭示其在准确、安全与多样性间的权衡。

Rethinking the Evaluation of Alignment Methods: Insights into Diversity, Generalisation, and Safety

  • 构建统一评估框架,从五个维度比较PPO、DPO等方法。
  • DPO和KTO最准,PPO最安全,PPO兼顾简洁与主动。
  • 适合想优化模型平衡性的研究人员参考。

大语言模型需在事实性、安全性、简洁性、主动性与多样性间精细对齐。现有研究多聚焦单一技术或特定维度,缺乏整体评估。本文提出统一评估框架,对比PPO、DPO、ORPO、KTO四种对齐方法在上述五维的表现,使用分布内与分布外数据集。通过经人类验证的LLM-as-Judge提示,发现DPO与KTO在事实准确性上领先,PPO与DPO在安全性上表现最优,而PPO在简洁性与主动性间取得最佳平衡。研究揭示了常见对齐方法的内在权衡,为开发更均衡可靠的模型提供指导。

原文摘要 · Abstract (English)

Large language models (LLMs) require careful alignment to balance competing objectives - factuality, safety, conciseness, proactivity, and diversity. Existing studies focus on individual techniques or specific dimensions, lacking a holistic assessment of the inherent trade-offs. We propose a unified evaluation framework that compares LLM alignment methods (PPO, DPO, ORPO, KTO) across these five axes, using both in-distribution and out-of-distribution datasets. Leveraging a specialized LLM-as-Judge prompt, validated through human studies, we reveal that DPO and KTO excel in factual accuracy, PPO and DPO lead in safety, and PPO best balances conciseness with proactivity. Our findings provide insights into trade-offs of common alignment methods, guiding the development of more balanced and reliable LLMs.

对齐评估大模型性能权衡

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。