arXiv:2604.12565cs.ROcs.CV2026-04

AutoMoMa高效生成大规模协同移动操作轨迹,突破数据瓶颈。

Scalable Trajectory Generation for Whole-Body Mobile Manipulation

  • 用统一运动链建模融合基座、机械臂与物体,实现并行优化
  • 每GPU小时生成5000条轨迹,超80倍快于传统方法
  • 适合研究真实复杂环境中机器人协同操作的学者

在非结构化环境中部署的机器人需协调全身运动——同时移动底盘和机械臂——以与物理世界交互。这种运动与灵巧性的耦合导致状态空间随场景和物体多样性呈组合增长,所需数据集远超固定基座操作。然而现有采集方法(如遥操作和规划)或劳动密集,或计算成本过高。核心瓶颈在于缺乏可扩展的管道来生成大规模、物理有效的协同轨迹数据。本文提出AutoMoMa,一个基于GPU加速的框架,将AKR建模(整合基座、机械臂与物体运动学为单一链)与并行轨迹优化相结合。AutoMoMa实现每GPU小时5000个轨迹(比基于CPU的基线快80倍以上),生成超过50万条物理有效的轨迹,覆盖330个场景、多样化刚性物体及多种机器人本体。先前数据集被迫在规模、多样性或运动学保真度间妥协;AutoMoMa同时解决三者。下游强化学习策略训练显示,即使单一带关节物体任务也需数万示范才能使最先进方法达到约80%成功率,证实数据稀缺才是关键约束。AutoMoMa因此连接高性能规划与可靠强化学习控制,填补了协同移动操作研究中长期缺失的数据基础设施。通过使大规模、运动学精确的训练数据成为可能,AutoMoMa展示了能在真实世界多样非结构化环境中通用运行的全身机器人策略。

原文摘要 · Abstract (English)

Robots deployed in unstructured environments must coordinate whole-body motion -- simultaneously moving a mobile base and arm -- to interact with the physical world. This coupled mobility and dexterity yields a state space that grows combinatorially with scene and object diversity, demanding datasets far larger than those sufficient for fixed-base manipulation. Yet existing acquisition methods, including teleoperation and planning, are either labor-intensive or computationally prohibitive at scale. The core bottleneck is the lack of a scalable pipeline for generating large-scale, physically valid, coordinated trajectory data across diverse embodiments and environments. Here we introduce AutoMoMa, a GPU-accelerated framework that unifies AKR modeling, which consolidates base, arm, and object kinematics into a single chain, with parallelized trajectory optimization. AutoMoMa achieves 5,000 episodes per GPU-hour (over $80\times$ faster than CPU-based baselines), producing a dataset of over 500k physically valid trajectories spanning 330 scenes, diverse articulated objects, and multiple robot embodiments. Prior datasets were forced to compromise on scale, diversity, or kinematic fidelity; AutoMoMa addresses all three simultaneously. Training downstream IL policies further reveals that even a single articulated-object task requires tens of thousands of demonstrations for SOTA methods to reach $\approx 80\%$ success, confirming that data scarcity -- not algorithmic limitations -- has been the binding constraint. AutoMoMa thus bridges high-performance planning and reliable IL-based control, providing the infrastructure previously missing for coordinated mobile manipulation research. By making large-scale, kinematically valid training data practical, AutoMoMa showcases generalizable whole-body robot policies capable of operating in the diverse, unstructured settings of the real world.

移动操作轨迹生成机器人学习大规模数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。