arXiv:2507.01961cs.ROcs.AI2025-07NeurIPS被引 27

提出AC-DiT模型,实现移动操作中基座与机械臂的自适应协调。

AC-DiT: Adaptive Coordination Diffusion Transformer for Mobile Manipulation

  • 通过基座运动条件机制,显式建模基座对机械臂的影响。
  • 动态调整2D图像与3D点云融合权重,适配不同阶段感知需求。
  • 在仿真与真实任务中均表现优越,适合复杂家居场景控制。

近期,移动操作因支持家庭任务中的语言控制而备受关注。然而,现有方法在协调移动基座与机械臂时仍面临挑战,主要源于两大局限:一方面,未显式建模基座运动对机械臂控制的影响,易在高自由度下导致误差累积;另一方面,全程采用单一视觉模态(如全2D或全3D),忽视了移动操作各阶段不同的多模态感知需求。为此,我们提出自适应协调扩散变压器(AC-DiT),增强端到端移动操作中基座与机械臂的协同。首先,由于基座运动直接影响机械臂动作,我们引入“基座到本体”条件机制,先提取基座运动表征,并作为预测全身动作的上下文先验,实现考虑基座运动影响的全身控制。其次,为满足不同阶段的感知需求,设计感知感知多模态条件策略,动态调整2D图像与3D点云的融合权重,生成符合当前感知需求的视觉特征。例如,在需要语义信息时更依赖2D输入,而在需精确空间理解时侧重3D几何信息。我们在仿真和真实世界移动操作任务中进行了大量实验验证,结果表明该方法显著提升性能。

原文摘要 · Abstract (English)

Recently, mobile manipulation has attracted increasing attention for enabling language-conditioned robotic control in household tasks. However, existing methods still face challenges in coordinating mobile base and manipulator, primarily due to two limitations. On the one hand, they fail to explicitly model the influence of the mobile base on manipulator control, which easily leads to error accumulation under high degrees of freedom. On the other hand, they treat the entire mobile manipulation process with the same visual observation modality (e.g., either all 2D or all 3D), overlooking the distinct multimodal perception requirements at different stages during mobile manipulation. To address this, we propose the Adaptive Coordination Diffusion Transformer (AC-DiT), which enhances mobile base and manipulator coordination for end-to-end mobile manipulation. First, since the motion of the mobile base directly influences the manipulator's actions, we introduce a mobility-to-body conditioning mechanism that guides the model to first extract base motion representations, which are then used as context prior for predicting whole-body actions. This enables whole-body control that accounts for the potential impact of the mobile base's motion. Second, to meet the perception requirements at different stages of mobile manipulation, we design a perception-aware multimodal conditioning strategy that dynamically adjusts the fusion weights between various 2D visual images and 3D point clouds, yielding visual features tailored to the current perceptual needs. This allows the model to, for example, adaptively rely more on 2D inputs when semantic information is crucial for action prediction, while placing greater emphasis on 3D geometric information when precise spatial understanding is required. We validate AC-DiT through extensive experiments on both simulated and real-world mobile manipulation tasks.

移动操作扩散模型多模态感知机器人控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。