让智能体通过符号化协商合作,避免动作错位冲突。
DR. WELL: Dynamic Reasoning and Learning with Symbolic World Model for Embodied LLM-Based Multi-Agent Collaboration
- 用符号世界模型进行分阶段协商,先定角色再生成计划。
- 在积木推移任务中,任务完成率与效率显著提升。
- 适合需要高可靠性协作的多智能体系统研究者。
协作式多智能体规划要求智能体在信息不全和通信受限下做出联合决策。轨迹级协调常因微小的时间或动作偏差而引发冲突。符号化规划通过提升抽象层级并提供最小化动作词汇表,实现同步与集体推进。我们提出 DR. WELL,一种去中心化的神经符号框架,用于协作式多智能体规划。协作通过两阶段协商协议展开:智能体首先基于推理提出候选角色,然后在共识与环境约束下达成联合分配。承诺后,各智能体独立生成并执行其角色的符号化计划,不暴露详细轨迹。计划通过共享世界模型与执行结果对齐,该模型编码当前状态并随智能体行动动态更新。通过在符号计划上推理而非原始轨迹,DR. WELL 避免了脆弱的步骤级对齐,支持可复用、可同步、可解释的高层操作。在协作积木推移任务上的实验表明,智能体能跨轮次适应,动态世界模型捕捉可复用模式,提升任务完成率与效率。实验显示,通过协商与自优化,动态世界模型以时间开销为代价,演化出更高效的协作策略。
原文摘要 · Abstract (English)
Cooperative multi-agent planning requires agents to make joint decisions with partial information and limited communication. Coordination at the trajectory level often fails, as small deviations in timing or movement cascade into conflicts. Symbolic planning mitigates this challenge by raising the level of abstraction and providing a minimal vocabulary of actions that enable synchronization and collective progress. We present DR. WELL, a decentralized neurosymbolic framework for cooperative multi-agent planning. Cooperation unfolds through a two-phase negotiation protocol: agents first propose candidate roles with reasoning and then commit to a joint allocation under consensus and environment constraints. After commitment, each agent independently generates and executes a symbolic plan for its role without revealing detailed trajectories. Plans are grounded in execution outcomes via a shared world model that encodes the current state and is updated as agents act. By reasoning over symbolic plans rather than raw trajectories, DR. WELL avoids brittle step-level alignment and enables higher-level operations that are reusable, synchronizable, and interpretable. Experiments on cooperative block-push tasks show that agents adapt across episodes, with the dynamic world model capturing reusable patterns and improving task completion rates and efficiency. Experiments on cooperative block-push tasks show that our dynamic world model improves task completion and efficiency through negotiation and self-refinement, trading a time overhead for evolving, more efficient collaboration strategies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。