系统梳理大模型对齐技术,揭示训练方法与安全机制的权衡关系。
Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges
- 对比监督微调与基于偏好优化的对齐方法,分析其适用场景差异。
- 指出当前评估框架存在奖励设定偏差、分布鲁棒性不足等关键缺陷。
- 适合关注大模型安全、价值对齐的研究者与工程师阅读。
由于大型语言模型(LLMs)展现出强大能力并日益影响社会各个层面,确保其与人类价值观和意图对齐已成为关键挑战。本文综述了大模型对齐的实际技术、训练范式及实证发现。我们分析了不同范式下对齐方法的发展,刻画了核心对齐目标间的根本权衡。研究表明,虽然监督微调可实现基础指令遵循,但基于偏好的方法更有利于对齐复杂的人类意图。我们讨论了前沿技术,包括直接偏好优化(DPO)、宪法式AI、类脑方法及对齐不确定性量化(AUQ),突出其在质量与效率间的平衡策略。我们回顾了现有评估框架与基准数据集,强调奖励设定偏差、分布鲁棒性及可扩展监督等局限性。总结了领先人工智能实验室采用的策略,以反映当前实践状态。最后,我们提出了监督、价值多元性、鲁棒性及持续对齐等方面的开放问题。本综述旨在为研究人员与从业者提供对大模型对齐演进格局的清晰指引。
原文摘要 · Abstract (English)
Due to the remarkable capabilities and growing impact of large language models (LLMs), they have been deeply integrated into many aspects of society. Thus, ensuring their alignment with human values and intentions has emerged as a critical challenge. This survey provides a comprehensive overview of practical alignment techniques, training protocols, and empirical findings in LLM alignment. We analyze the development of alignment methods across diverse paradigms, characterizing the fundamental trade-offs between core alignment objectives. Our analysis shows that while supervised fine-tuning enables basic instruction-following, preference-based methods offer more flexibility for aligning with nuanced human intent. We discuss state-of-the-art techniques, including Direct Preference Optimization (DPO), Constitutional AI, brain-inspired methods, and alignment uncertainty quantification (AUQ), highlighting their approaches to balancing quality and efficiency. We review existing evaluation frameworks and benchmarking datasets, emphasizing limitations such as reward misspecification, distributional robustness, and scalable oversight. We summarize strategies adopted by leading AI labs to illustrate the current state of practice. We conclude by outlining open problems in oversight, value pluralism, robustness, and continuous alignment. This survey aims to inform both researchers and practitioners navigating the evolving landscape of LLM alignment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。