arXiv:2512.08545cs.CLcs.AI2025-12被引 1

64×64智能体网格分阶段解决长时序任务,提升稳定性与效率。

Curriculum Guided Massive Multi Agent System Solving For Robust Long Horizon Tasks

  • 构建64×64轻量级智能体网格,按空间课程逐步扩展任务范围。
  • 通过负对数似然评估置信度,降低对专家反馈依赖37%以上。
  • 适合需要长期规划的机器人操控与复杂任务分解场景。

大型语言模型与多智能体系统在分解复杂任务方面展现出潜力,但在长时序推理任务中表现不佳且计算成本持续上升。本文提出一种分层多智能体架构,将推理分布于64×64的轻量级智能体网格上,并由选择性专家支持。通过空间课程逐步扩大网格的运行区域,确保智能体先掌握中心区域的简单任务,再逐步应对外围更复杂的任务。为提升可靠性,系统引入负对数似然(NLL)作为置信度度量,使课程能优先选择既准确又校准良好的区域。采用汤普森采样课程管理器,根据智能体能力与基于NLL的奖励信号自适应选择训练区域。在具有空间结构的汉诺塔基准上进行评估,该任务模拟了诸多机器人操作与规划任务的长时序特性。结果表明,该方法显著提升了系统稳定性,减少超过37%的专家调用次数,并通过分布式协作实现更强的长程推理能力。

原文摘要 · Abstract (English)

Large Language Models and multi-agent systems have shown promise in decomposing complex tasks, yet they struggle with long-horizon reasoning tasks and escalating computation cost. This work introduces a hierarchical multi-agent architecture that distributes reasoning across a 64*64 grid of lightweight agents, supported by a selective oracle. A spatial curriculum progressively expands the operational region of the grid, ensuring that agents master easier central tasks before tackling harder peripheral ones. To improve reliability, the system integrates Negative Log-Likelihood as a measure of confidence, allowing the curriculum to prioritize regions where agents are both accurate and well calibrated. A Thompson Sampling curriculum manager adaptively chooses training zones based on competence and NLL-driven reward signals. We evaluate the approach on a spatially grounded Tower of Hanoi benchmark, which mirrors the long-horizon structure of many robotic manipulation and planning tasks. Results demonstrate improved stability, reduced oracle usage, and stronger long-range reasoning from distributed agent cooperation.

多智能体长时序推理课程学习机器人规划

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。