arXiv:2505.20726cs.RO2025-05被引 4

自动生成任意场景下的多样化机器人任务,用于评估和提升智能体决策能力。

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making

  • 基于场景自动生成过程型与目标型任务,覆盖广泛可行性。
  • 在仿真与真实场景中验证任务有效性,支持多智能体评测。
  • 可直接用于构建基准与优化视觉语言模型的决策能力。

构建能够完成任意任务的具身智能体是实现具身通用人工智能(E-AGI)的核心目标。尽管近期研究推进了通用机器人策略的发展,但其训练与评估通常局限于特定场景、有限指令与情境。现有基准也多依赖人工标注的少量任务。我们认为,探索给定场景内所有可行任务至关重要,既能提供丰富评测基准,也能为智能体改进提供资源。为此,我们提出 ManiTaskGen,一个针对任意场景的自动任务生成系统,可生成涵盖过程型(如“将物体从X移动到Y”)与结果型(如“清理桌面”)指令的多样化、可执行任务。我们在仿真与真实场景中应用该系统,验证了生成任务的有效性与多样性。进一步利用这些任务自动构建基准,全面评估基于现有视觉语言模型(VLMs)的具身决策能力。此外,我们提出一种简单有效的增强方法,利用生成任务提升智能体决策表现。本工作建立了一个适用于任意场景的通用任务生成框架,推动具身决策智能体的评测与优化。

原文摘要 · Abstract (English)

Building embodied agents capable of accomplishing arbitrary tasks is a core objective towards achieving embodied artificial general intelligence (E-AGI). While recent work has advanced such general robot policies, their training and evaluation are often limited to tasks within specific scenes, involving restricted instructions and scenarios. Existing benchmarks also typically rely on manual annotation of limited tasks in a few scenes. We argue that exploring the full spectrum of feasible tasks within any given scene is crucial, as they provide both extensive benchmarks for evaluation and valuable resources for agent improvement. Towards this end, we introduce ManiTaskGen, a novel system that automatically generates comprehensive, diverse, feasible mobile manipulation tasks for any given scene. The generated tasks encompass both process-based, specific instructions (e.g., "move object from X to Y") and outcome-based, abstract instructions (e.g., "clear the table"). We apply ManiTaskGen to both simulated and real-world scenes, demonstrating the validity and diversity of the generated tasks. We then leverage these tasks to automatically construct benchmarks, thoroughly evaluating the embodied decision-making capabilities of agents built upon existing vision-language models (VLMs). Furthermore, we propose a simple yet effective method that utilizes ManiTaskGen tasks to enhance embodied decision-making. Overall, this work presents a universal task generation framework for arbitrary scenes, facilitating both benchmarking and improvement of embodied decision-making agents.

具身智能任务生成视觉语言模型机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。