arXiv:2412.11974cs.ROcs.AI2024-12被引 47

让机器人像人一样思考空间动作,实现复杂任务的精准执行

Emma-X: An Embodied Multimodal Action Model with Grounded Chain of Thought and Look-ahead Spatial Reasoning

  • 构建分层实体化数据集,让模型学会基于真实动作推理任务步骤
  • 在6万条机器人操作轨迹上训练,实测表现优于现有方法
  • 适合研究具身智能、机器人长程规划与多模态决策的学者

传统强化学习机器人控制方法通常任务特定,难以跨环境或面对新物体、新指令时泛化。视觉语言模型(VLM)虽具备强场景理解与规划能力,但无法生成适配具体机器人形态的动作策略。为此,视觉-语言-动作(VLA)模型应运而生,但仍面临长程空间推理与可落地任务规划的挑战。本文提出具身多模态动作模型Emma-X,基于BridgeV2构建包含6万条机器人操作轨迹的分层实体化数据集,并自动标注了具身任务推理与空间引导信息。我们还设计了一种基于夹爪状态与运动轨迹的轨迹分割策略,有效缓解子任务推理生成中的幻觉问题。实验表明,Emma-X在需要空间推理的真实机器人任务中显著优于对比基线。

原文摘要 · Abstract (English)

Traditional reinforcement learning-based robotic control methods are often task-specific and fail to generalize across diverse environments or unseen objects and instructions. Visual Language Models (VLMs) demonstrate strong scene understanding and planning capabilities but lack the ability to generate actionable policies tailored to specific robotic embodiments. To address this, Visual-Language-Action (VLA) models have emerged, yet they face challenges in long-horizon spatial reasoning and grounded task planning. In this work, we propose the Embodied Multimodal Action Model with Grounded Chain of Thought and Look-ahead Spatial Reasoning, Emma-X. Emma-X leverages our constructed hierarchical embodiment dataset based on BridgeV2, containing 60,000 robot manipulation trajectories auto-annotated with grounded task reasoning and spatial guidance. Additionally, we introduce a trajectory segmentation strategy based on gripper states and motion trajectories, which can help mitigate hallucination in grounding subtask reasoning generation. Experimental results demonstrate that Emma-X achieves superior performance over competitive baselines, particularly in real-world robotic tasks requiring spatial reasoning.

具身智能多模态机器人控制空间推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。