arXiv:2501.17112cs.LG2025-01被引 3

改进逆宪法人工智能算法,让大模型对齐更透明可解释。

Decoding Human Preferences in Alignment: An Improved Approach to Inverse Constitutional AI

  • 通过优化原则生成与聚类,从偏好数据中提取更准确的规则
  • 在合成与真实数据集上均提升原则的准确性和泛化能力
  • 适合关注模型对齐透明性与可解释性的研究人员

传统的大语言模型对齐方法如基于人类反馈的强化学习(RLHF)和直接偏好优化(DPO)依赖隐式原则,限制了可解释性。宪法人工智能(CAI)提供了一种显式的规则驱动框架来指导对齐。在此基础上,我们改进了逆宪法人工智能(ICAI)算法,该算法从偏好数据集中提取宪法规则。通过优化原则生成、聚类及嵌入过程,新方法在合成与真实世界数据集上均提升了所提取原则的准确性与泛化能力。结果表明,这些规则有望推动更透明、更可适应的对齐方法发展,为超越传统微调的未来方向提供了可能。

原文摘要 · Abstract (English)

Traditional methods for aligning Large Language Models (LLMs), such as Reinforcement Learning from Human Feedback (RLHF) and Direct Preference Optimization (DPO), rely on implicit principles, limiting interpretability. Constitutional AI (CAI) offers an explicit, rule-based framework for guiding LLM alignment. Building on this, we refine the Inverse Constitutional AI (ICAI) algorithm, which extracts constitutions from preference datasets. By improving principle generation, clustering, and embedding processes, our approach enhances the accuracy and generalizability of extracted principles across synthetic and real-world datasets. Our results highlight the potential of these principles to foster more transparent and adaptable alignment methods, offering a promising direction for future advancements beyond traditional fine-tuning.

模型对齐可解释性偏好学习规则生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。