arXiv:2605.12240cs.AI2026-05

通过分层协作机制提升长任务服务代理的可靠性

No Action Without a NOD: A Heterogeneous Multi-Agent Architecture for Reliable Service Agents

论文配图:No Action Without a NOD: A Heterogeneous Multi-Agent Architecture for Reliable Service Agents
图 1 · 摘自论文原文
  • 引入导航-操作-监督三角色架构,显式追踪任务状态
  • 在关键动作前增加独立审核,降低错误传播风险
  • 显著减少策略违规和工具幻觉,适合高可靠场景

大型语言模型代理在预订机票等服务应用中日益普及,但在长周期任务中常出现策略违规、工具幻觉和行为偏差,严重阻碍其实际部署。为此,我们提出NOD(Navigator-Operator-Director)异构多智能体架构。不同于以往隐式依赖对话上下文维护任务状态的方法,NOD将结构化全局状态外部化,使导航者能够显式追踪并一致决策。同时,在关键动作前引入选择性外部监督,由独立的指挥者代理进行执行验证并在必要时干预。实验表明,在τ²-Bench数据集上,NOD相比基线方法提升了任务成功率与关键动作精确率,并显著降低了策略违规、工具幻觉和用户意图错位问题。

原文摘要 · Abstract (English)

Large language model (LLM) agents have increasingly advanced service applications, such as booking flight tickets. However, these service agents suffer from unreliability in long-horizon tasks, as they often produce policy violations, tool hallucinations, and misaligned actions, which greatly impedes their real-world deployment. To address these challenges, we propose NOD (Navigator-Operator-Director), a heterogeneous multi-agent architecture for service agents. Instead of maintaining task state implicitly in dialogue context as in prior work, we externalize a structured Global State to enable explicit task state tracking and consistent decision-making by the Navigator. Besides, we introduce selective external oversight before critical actions, allowing an independent Director agent to verify execution and intervene when necessary. As such, NOD effectively mitigates error propagation and unsafe behavior in long-horizon tasks. Experiments on $τ^2$-Bench demonstrate that NOD achieves higher task success rates and critical action precision over baselines. More importantly, NOD improves the reliability of service agents by reducing policy violations, tool hallucinations, and user-intent misalignment.

多智能体服务代理可靠性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。