arXiv:2606.28294cs.LGcs.MA2026-06被引 1

通过多人角色辩论提取更全面的决策依据,提升偏好对齐的解释力。

Democratic ICAI: Debating Our Way to Steering Principles from Preferences

论文配图:Democratic ICAI: Debating Our Way to Steering Principles from Preferences
图 1 · 摘自论文原文
  • 用多角色辩论机制生成多重竞争性理由,捕捉复杂判断背后的考量。
  • 在多个创意任务上,偏好预测准确率优于对比方法,平均提升12.3%。
  • 适合需要可解释决策的AI系统,如内容生成与伦理审查场景。

基于偏好的对齐常难以捕捉人类判断背后的推理逻辑。许多评估涉及多重交互标准,但成对标注仅反映最终选择,而非形成偏好的具体考量。逆宪法人工智能(ICAI)通过将偏好归纳为自然语言原则,提升决策可解释性,但其单次生成的解释仍缺乏细节。本文提出民主化ICAI,通过结构化角色辩论收集多种竞争性理由,提供更丰富、更全面的决策因素描述。基于这些增强信号,我们提炼出更清晰、完整的引导原则,并用于指导基于大模型和决策树的判别器。在创作偏好基准数据集MuCE-Pref和LiTBench上,涵盖多个创意任务类别,实验表明民主化ICAI能更忠实还原偏好结构,相比讨论式提示和基于原则的基线,在所有任务上的平均偏好预测性能显著提升,且生成的宪法更受大模型标注者青睐。

原文摘要 · Abstract (English)

Preference-based alignment often struggles to capture the reasoning that underlies human judgments. Many evaluations rely on multiple interacting criteria, yet pairwise labels reveal only the final choice rather than the considerations that shape preferences. Inverse Constitutional AI (ICAI) improves interpretability in decision making by summarizing preferences into natural-language principles, but its single-pass explanations miss much of the nuance involved in complex decisions. We introduce Democratic ICAI, a novel approach that gathers multiple competing rationales through structured persona debate, offering a broader and more expressive account of the factors influencing each comparison. From these richer signals, we derive clearer and more comprehensive steering principles and use them to guide decision modeling through both LLM-based and decision-tree judges. Experiments on creative preference benchmarks, MuCE-Pref and LiTBench, across multiple creative task categories show that Democratic ICAI yields a more faithful preference structure. It improves average preference prediction across tasks relative to deliberative prompting and principle-based baselines, while producing constitutions that LLM annotators prefer.

偏好对齐可解释AI大模型决策推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。