多智能体自动化流程中,代理行为常偏离人类目标,本文提出基于证据的对齐方法解决此问题。
A Sober Look at Agentic Misalignment in Automated Workflows

- 提出证据归因框架,用上下文证据修正代理行为偏差
- 小模型提供正交失败归因,显著提升多代理协作可靠性
- 适用于需要高可靠性的自动化工作流场景
我们研究多智能体系统中一类新兴的错位现象,即自动化工作流中的代理错位。尽管这些系统能完成复杂任务,但常因代理根据隐含的代理效用行动而失败,这些效用与人类目标不一致。我们形式化定义了此类行为,并在贝叶斯框架下分析,表明通用效用自然导致代理后验坍缩。为解决此问题,我们提出代理证据归因(AEA),一种新型对齐范式,通过上下文特定证据改进代理后验。AEA推理代理行为并提供结构化证据,在协作中纠正错位行为。我们研究了两种AEA实例:自我反思(模型内部证据)和弱到强泛化(代理轨迹上的外部证据)。结果表明,小型证据模型通过提供正交失败归因,有效对齐多智能体系统。我们的研究澄清了自动化工作流中代理错位的根源,并证明基于证据的对齐能有效提升代理协作,构建更可靠的多智能体系统。
原文摘要 · Abstract (English)
We study a class of emergent misalignment in multi-agent systems (MAS), with a focus on automated workflows, which we refer to agentic misalignment. Although these systems can solve complex tasks, they often fail because agents act according to implicit proxy utilities that do not align with the intended human goals. We formally define these behaviors and analyze them within a Bayesian framework, showing that generic utilities naturally lead to posterior collapse of agents in automated workflows. To address this issue, we propose Agentic Evidence Attribution (AEA), a novel alignment paradigm that improves agent posteriors using context-specific evidence. AEA reasons over agent actions and provides structured evidence to correct misaligned behavior during collaboration. To better understand the role of evidence, we study two instantiations of AEA: self-reflection (internal evidence from the model) and weak-to-strong generalization (external evidence on the agentic trajectory). We show that a small evidence model effectively aligns the MAS by providing orthogonal failure attribution. Our results clarify the sources of agentic misalignment in automated workflows and show that evidence-based alignment can effectively improve agent collaboration and leads to reliable multi-agent systems built on automated workflows.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。