arXiv:2606.30116cs.AI2026-06

用自然语言原则重构模型偏好,发现执行者差异导致结果不一致。

Open Problems in Constitutional Preference Reconstruction

  • 将偏好数据压缩为可解释的自然语言原则,但组合规则未明确定义。
  • 不同执行者对同一套原则判断一致率仅73%,跨模型一致率也仅73%。
  • 改进版ICAI+提升执行一致性,适合关注模型可解释性的研究者。

成对偏好数据广泛用于训练和评估语言模型(如RLHF),但每个数据点仅记录选择结果,未包含其背后的理由。逆宪法人工智能(ICAI)等方法试图通过将数据集压缩为简短的自然语言原则“宪法”来提升可解释性。我们指出这一框架尚不明确:单一原则列表无法构成可执行的决策规则,因原则间的组合方式仍隐含未明。本文以成对设定为测试平台,实证揭示宪法方法中的三个开放问题:第一,原则质量难以衡量,覆盖率与准确率仅为端到端重构的不完整代理;第二,原则组合存在歧义:在固定原则下,不同执行者(大模型裁判与多数投票)间一致率仅为73%;第三,不同大模型生成的宪法差异显著:跨模型投票一致率为73%,而同一模型内部一致率达81%。在PRISM、AlpacaEval与Chatbot Arena上,我们发现原则优化(ICAI+)或可缓解这些问题:跨执行者一致率升至78%,透明执行器准确率匹配大模型裁判(66% vs. 67%)。结果表明,宪法应作为“宪法-执行系统”整体评估,对大模型作为裁判的应用具有普遍意义。

原文摘要 · Abstract (English)

Pairwise preference data is widely used for training and evaluating language models (e.g., RLHF), but each datapoint records a \emph{choice}, not the rationale behind it. Methods such as Inverse Constitutional AI (ICAI) attempt to improve interpretability by compressing datasets into short ``constitutions'' of natural-language principles. We argue this framing is under-specified: a flat list of principles is not yet an executable decision rule because it leaves principle composition implicit. We use the pairwise setting as a testbed to empirically characterize three open problems in constitutional methods. First, principle quality is hard to measure: coverage and accuracy are useful but incomplete proxies for end-to-end reconstruction. Second, \emph{composition is ambiguous}: holding principles fixed, different executors (LLM judge versus majority vote) agree only $73\%$ of the time. Third, \emph{constitutions differ between LLMs}: cross-model vote agreement is $73\%$, whereas intra-model agreement is $81\%$. Across PRISM, AlpacaEval, and Chatbot Arena, we show that principle refinement (ICAI+) may be a first step towards ameliorating these problems: inter-executor agreement rises to $78\%$, and transparent executors match LLM judge accuracy ($66\%$ vs.\ $67\%$). Our results highlight that constitutions should be evaluated as \emph{constitution--executor systems}, with implications for LLMs-as-a-judge broadly.

偏好学习可解释性大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。