arXiv:2506.10966cs.RO2025-06CVPR被引 21

用大模型生成多样化任务,测试机器人指令泛化能力

GENMANIP: LLM-driven Simulation for Generalizable Instruction-Following Manipulation

  • 用大模型自动生成带场景图的复杂任务
  • 基于200个人工校正场景,发现模块化系统更泛化
  • 适合研究机器人指令理解与跨场景适应

现实世界中的机器人操作仍面临鲁棒泛化挑战。现有仿真平台难以支持策略在多样指令和场景下的适应性研究,滞后于指令跟随基础模型(如大语言模型)的发展。为此,我们提出GenManip——一个面向政策泛化研究的逼真桌面仿真平台,通过大模型驱动的任务导向场景图,利用10,000个标注3D物体资产自动生成大规模、多样化的任务。为系统评估泛化能力,我们构建了经人机协同修正的GenManip-Bench基准,包含200个场景。评估两种策略:(1) 集成基础模型的模块化系统(感知、推理、规划),(2) 通过可扩展数据收集训练的端到端策略。结果表明,尽管数据规模提升对端到端方法有益,但融合基础模型的模块化系统在多场景中表现更优。该平台有望推动真实条件下政策泛化的研究。

原文摘要 · Abstract (English)

Robotic manipulation in real-world settings remains challenging, especially regarding robust generalization. Existing simulation platforms lack sufficient support for exploring how policies adapt to varied instructions and scenarios. Thus, they lag behind the growing interest in instruction-following foundation models like LLMs, whose adaptability is crucial yet remains underexplored in fair comparisons. To bridge this gap, we introduce GenManip, a realistic tabletop simulation platform tailored for policy generalization studies. It features an automatic pipeline via LLM-driven task-oriented scene graph to synthesize large-scale, diverse tasks using 10K annotated 3D object assets. To systematically assess generalization, we present GenManip-Bench, a benchmark of 200 scenarios refined via human-in-the-loop corrections. We evaluate two policy types: (1) modular manipulation systems integrating foundation models for perception, reasoning, and planning, and (2) end-to-end policies trained through scalable data collection. Results show that while data scaling benefits end-to-end methods, modular systems enhanced with foundation models generalize more effectively across diverse scenarios. We anticipate this platform to facilitate critical insights for advancing policy generalization in realistic conditions. Project Page: https://genmanip.axi404.top/.

机器人操作大模型仿真平台泛化能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。