arXiv:2509.09332cs.ROcs.AI2025-09被引 4

让机器人计划更智能:能根据任务自动选3D信息,还考虑实际物理限制。

OmniEVA: Embodied Versatile Planner via Task-Adaptive 3D-Grounded and Embodiment-aware Reasoning

  • 用动态路由选择性融合3D信息,按任务需求灵活适配空间理解。
  • 在10个基准任务上表现领先,复合任务成功率超90%。
  • 适合研究具身智能、机器人规划的学者和开发者。

多模态大语言模型(MLLM)为具身智能带来新可能,支持多模态理解、推理与连续空间决策。但现有系统存在两大瓶颈:一是几何适应性差距——仅基于2D输入或硬编码3D结构的模型,要么缺乏空间信息,要么泛化能力差;二是具身约束差距——忽略真实机器人的物理限制,导致计划理论上可行却无法执行。为此,我们提出OmniEVA:一种通过任务自适应3D对齐与具身感知推理实现通用具身规划的新框架。其核心创新包括:(1) 任务自适应3D接地机制,通过门控路由器根据上下文需求显式调控3D融合,实现对多样化具身任务的上下文感知3D对齐;(2) 具身感知推理框架,将任务目标与具身约束联合纳入推理循环,生成既目标导向又可执行的规划。大量实验表明,OmniEVA不仅在多项具身推理任务中达到领先性能,且在涵盖基础与复合任务的系列新基准上展现出强鲁棒性与泛化能力。

原文摘要 · Abstract (English)

Recent advances in multimodal large language models (MLLMs) have opened new opportunities for embodied intelligence, enabling multimodal understanding, reasoning, and interaction, as well as continuous spatial decision-making. Nevertheless, current MLLM-based embodied systems face two critical limitations. First, Geometric Adaptability Gap: models trained solely on 2D inputs or with hard-coded 3D geometry injection suffer from either insufficient spatial information or restricted 2D generalization, leading to poor adaptability across tasks with diverse spatial demands. Second, Embodiment Constraint Gap: prior work often neglects the physical constraints and capacities of real robots, resulting in task plans that are theoretically valid but practically infeasible. To address these gaps, we introduce OmniEVA -- an embodied versatile planner that enables advanced embodied reasoning and task planning through two pivotal innovations: (1) a Task-Adaptive 3D Grounding mechanism, which introduces a gated router to perform explicit selective regulation of 3D fusion based on contextual requirements, enabling context-aware 3D grounding for diverse embodied tasks. (2) an Embodiment-Aware Reasoning framework that jointly incorporates task goals and embodiment constraints into the reasoning loop, resulting in planning decisions that are both goal-directed and executable. Extensive experimental results demonstrate that OmniEVA not only achieves state-of-the-art general embodied reasoning performance, but also exhibits a strong ability across a wide range of downstream scenarios. Evaluations of a suite of proposed embodied benchmarks, including both primitive and composite tasks, confirm its robust and versatile planning capabilities. Project page: https://omnieva.github.io

具身智能任务规划3D推理机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。