arXiv:2512.06196cs.AIcs.CL2025-12中稿 · AAAI被引 3

让大模型代理按需调整行为,无需重训即可改变偏好。

ARCANE: A Multi-Agent Framework for Interpretable and Configurable Alignment

  • 用自然语言规则动态表示用户偏好,支持实时修改。
  • 在219个标注规则上测试,实现正确性与简洁性的可配置平衡。
  • 适合需要透明、灵活调整的复杂长任务系统使用。

基于大语言模型的智能体在执行长期任务时,保持与利益相关方偏好一致至关重要。有效对齐需要可解释的奖励模型,使利益相关方能够理解并审计模型目标;同时,奖励模型需能在交互时引导智能体,实现偏好调整而无需重新训练。我们提出ARCANE框架,将对齐视为多智能体协作问题,将利益相关方偏好动态建模为自然语言规则:可验证的标准加权集合,能根据任务上下文实时生成。受效用理论启发,我们将规则学习建模为重构问题,并采用正则化的分组序列策略优化(GSPO)方法,在可解释性、忠实性和计算效率间取得平衡。基于219个来自GDPVal基准的标注规则语料,我们在需要多步推理和工具使用的挑战性任务上评估了ARCANE。学习得到的规则生成了紧凑、易读的评估结果,并实现了无需重训的可配置权衡(如正确性与简洁性)。结果表明,基于规则的奖励模型为复杂、长期的AI系统提供了可解释、运行时自适应对齐的可行路径。

原文摘要 · Abstract (English)

As agents based on large language models are increasingly deployed to long-horizon tasks, maintaining their alignment with stakeholder preferences becomes critical. Effective alignment in such settings requires reward models that are interpretable so that stakeholders can understand and audit model objectives. Moreover, reward models must be capable of steering agents at interaction time, allowing preference shifts to be incorporated without retraining. We introduce ARCANE, a framework that frames alignment as a multi-agent collaboration problem that dynamically represents stakeholder preferences as natural-language rubrics: weighted sets of verifiable criteria that can be generated on-the-fly from task context. Inspired by utility theory, we formulate rubric learning as a reconstruction problem and apply a regularized Group-Sequence Policy Optimization (GSPO) procedure that balances interpretability, faithfulness, and computational efficiency. Using a corpus of 219 labeled rubrics derived from the GDPVal benchmark, we evaluate ARCANE on challenging tasks requiring multi-step reasoning and tool use. The learned rubrics produce compact, legible evaluations and enable configurable trade-offs (e.g., correctness vs. conciseness) without retraining. Our results show that rubric-based reward models offer a promising path toward interpretable, test-time adaptive alignment for complex, long-horizon AI systems.

智能体对齐可解释性动态偏好多智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。