arXiv:2512.24653cs.RO2025-12被引 27

构建31万条双臂机器人操作数据集,支持复杂任务泛化与真实世界迁移。

RoboMIND 2.0: A Multimodal, Bimanual Mobile Manipulation Dataset for Generalizable Embodied Intelligence

  • 采集6种机器人在739项任务上的31万条双臂操作轨迹,含触觉增强与移动操作数据。
  • 提出MIND-2系统,结合语义规划与视觉-语言-动作执行,实现自然语言到动作的精准映射。
  • 配套20万条仿真数据,助力模拟到真实的鲁棒迁移,适合具身智能研究者使用。

尽管基于数据的模仿学习已革新机器人操作,但现有方法仍受限于大规模、多样化真实世界示范数据的缺乏,导致模型在长时序双臂任务及非结构化环境中的移动操作泛化能力有限。为弥合这一差距,我们推出了RoboMIND 2.0,一个综合性的真实世界数据集,包含超过31万条双臂操作轨迹,覆盖六种不同机器人形态和739个复杂任务。关键的是,为支持高接触频率和空间扩展任务的研究,数据集整合了1.2万条触觉增强样本和2万条移动操作轨迹。此外,我们构建了真实环境的高保真数字孪生,发布额外2万条仿真轨迹数据,以促进稳健的模拟到现实迁移。为充分挖掘RoboMIND 2.0潜力,我们提出MIND-2系统,一种通过离线强化学习优化的分层双系统框架。MIND-2集成高层语义规划器(MIND-2-VLM),将抽象自然语言指令分解为具象子目标,并结合低层视觉-语言-动作执行器(MIND-2-VLA),生成精确且具备本体感知的运动动作。

原文摘要 · Abstract (English)

While data-driven imitation learning has revolutionized robotic manipulation, current approaches remain constrained by the scarcity of large-scale, diverse real-world demonstrations. Consequently, the ability of existing models to generalize across long-horizon bimanual tasks and mobile manipulation in unstructured environments remains limited. To bridge this gap, we present RoboMIND 2.0, a comprehensive real-world dataset comprising over 310K dual-arm manipulation trajectories collected across six distinct robot embodiments and 739 complex tasks. Crucially, to support research in contact-rich and spatially extended tasks, the dataset incorporates 12K tactile-enhanced episodes and 20K mobile manipulation trajectories. Complementing this physical data, we construct high-fidelity digital twins of our real-world environments, releasing an additional 20K-trajectory simulated dataset to facilitate robust sim-to-real transfer. To fully exploit the potential of RoboMIND 2.0, we propose MIND-2 system, a hierarchical dual-system frame-work optimized via offline reinforcement learning. MIND-2 integrates a high-level semantic planner (MIND-2-VLM) to decompose abstract natural language instructions into grounded subgoals, coupled with a low-level Vision-Language-Action executor (MIND-2-VLA), which generates precise, proprioception-aware motor actions.

机器人操作多模态数据具身智能仿真迁移

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。