arXiv:2510.11660cs.ROcs.AI2025-10被引 6

ManiAgent让机器人通过多智能体协作完成复杂抓取任务,成功率超95%。

ManiAgent: An Agentic Framework for General Robotic Manipulation

  • 多智能体协作分工:感知、拆解任务、生成动作
  • 真实场景抓放任务成功率95.8%,模拟环境达86.8%
  • 可高效收集数据,训练出媲美人工标注的视觉语言模型

尽管视觉-语言-动作(VLA)模型在机器人操作中表现优异,但在复杂推理和长程任务规划方面仍受限于数据稀缺与模型容量。为此,我们提出ManiAgent,一种面向通用操作任务的智能体架构,实现从任务描述和环境输入到机器人操作动作的端到端输出。该框架中,多个智能体通过协同通信完成环境感知、子任务分解与动作生成,有效应对复杂操作场景。评估显示,ManiAgent在SimplerEnv基准上达到86.8%的成功率,在真实世界的抓放任务中达95.8%,并能高效生成数据,使VLA模型性能接近基于人工标注数据训练的水平。

原文摘要 · Abstract (English)

While Vision-Language-Action (VLA) models have demonstrated impressive capabilities in robotic manipulation, their performance in complex reasoning and long-horizon task planning is limited by data scarcity and model capacity. To address this, we introduce ManiAgent, an agentic architecture for general manipulation tasks that achieves end-to-end output from task descriptions and environmental inputs to robotic manipulation actions. In this framework, multiple agents involve inter-agent communication to perform environmental perception, sub-task decomposition and action generation, enabling efficient handling of complex manipulation scenarios. Evaluations show ManiAgent achieves an 86.8% success rate on the SimplerEnv benchmark and 95.8% on real-world pick-and-place tasks, enabling efficient data collection that yields VLA models with performance comparable to those trained on human-annotated datasets. The project webpage is available at https://yi-yang929.github.io/ManiAgent/.

机器人操作多智能体端到端

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。