arXiv:2601.18760cs.LGcs.CL2026-01被引 2

让AI原则基于真实人类理由与价值观,提升对齐的公平性与可信度。

Beyond Preferences: Learning Alignment Principles Grounded in Human Reasons and Values

  • 结合用户交互时的理由与深层价值观生成AI行为准则
  • 新准则在偏好测试中胜过传统方法,更受人类认可
  • 适合关注AI伦理、政策制定及可信对齐的研究者

大语言模型的对齐需考虑人类价值观。当前宪法式对齐依赖自然语言描述的原则,但如何公平地整合多方意见尚不明确。本文提出基于理由与价值观的宪法生成框架(GCAI),通过分析用户对偏好提供的理由,生成情境化原则;同时从用户表达的对AI的价值观中提取通用原则。实验表明,由GCAI生成的宪法在个人偏好和广泛适用性上均优于传统逆向宪法对齐方法(ICAI),且被评价为更具道德基础、逻辑一致性和多元包容性。

原文摘要 · Abstract (English)

A crucial consideration when developing and deploying Large Language Models (LLMs) is the human values to which these models are aligned. In the constitutional framework of alignment models are aligned to a set of principles (the constitution) specified in natural language. However, it is unclear how to fairly determine this constitution with widespread stakeholder input. In this work we propose Grounded Constitutional AI (GCAI), a unified framework for generating constitutions of principles that are representative of both users' general expectations toward AI (general principles) and their interaction-time preferences (contextual principles). We extend the Inverse Constitutional AI (ICAI) approach to generate contextual principles from human preference annotation data by leveraging human-provided \textit{reasons} for their preferences. We supplement these contextual principles with general principles surfaced from user statements of \textit{values} regarding AI. We show that a constitution generated by GCAI is preferred by humans over one generated through ICAI both personally, and for widespread use in governing AI behavior. Additionally participants consider the GCAI constitution to be more morally grounded, coherent, and pluralistic.

AI对齐价值观对齐伦理准则

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。