arXiv:2605.23414cs.AIcs.LG2026-05

大模型多智能体系统常因认知误判失败,新方法提升计划可靠性

When Planning Fails Despite Correct Execution: On Epistemic Calibration for LLM-Based Multi-Agent Systems

论文配图:When Planning Fails Despite Correct Execution: On Epistemic Calibration for LLM-Based Multi-Agent Systems
图 1 · 摘自论文原文
  • 通过信息一致性评估计划稳定性,避免认知误判
  • 实验显示系统成功率平均提升9.75%
  • 适合需要长期可靠决策的复杂协作场景

基于大语言模型的多智能体系统即使执行正确仍会失败,原因在于智能体在评估计划可行性时会误判自身知识水平,这种现象称为规划中的认知误校准。与执行错误不同,认知误校准在规划阶段是隐性的,生成的计划看似自洽且可执行,无明显错误;同时该误差具有动态性,新信息可能改变可行性判断,掩盖过去的误校准信号,导致问题反复出现。为此,我们提出认知规划校准智能体工作流(EPC-AW),不直接验证可行性,而是评估计划在不同信息条件下是否依然成立。EPC-AW采用基于信息一致性的计划选择机制,挑选在多个智能体间评估结果稳定的计划,并结合一致性引导的认知状态修正策略,利用历史偏差指导未来规划。实验表明,该方法使系统整体成功率平均提升9.75%。

原文摘要 · Abstract (English)

LLM-based multi-agent systems can fail even when planned actions are executed correctly because agents may misjudge their knowledge when evaluating plan feasibility, a phenomenon we term epistemic miscalibration in planning. Unlike execution errors, epistemic miscalibration is latent during planning, as generated plans can remain self-consistent and executable without observable errors; the miscalibration is also dynamic, as new information can alter feasibility assessments, potentially obscuring past miscalibration signals and causing them to recur over time. To address this, we propose the Epistemic Planning Calibration Agentic Workflow (EPC-AW), which assesses whether plans remain supported under varying information conditions rather than directly verifying feasibility. EPC-AW employs Information-consistency-based Plan Selection, selecting plans whose evaluations are stable across agents, together with Consistency-guided Epistemic State Refinement to adapt calibration over time by leveraging past discrepancies to guide future planning. Experiments show that EPC-AW improves system-level success by an average of 9.75%.

大模型多智能体认知校准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。