提出DSPv2架构,提升机器人全身操作的泛化与精准控制能力
DSPv2: Improved Dense Policy for Effective and Generalizable Whole-body Mobile Manipulation
- 融合3D空间特征与多视角2D语义特征,增强感知理解
- 在多个任务上显著优于现有方法,泛化性能突出
- 适合需要复杂环境适应的机器人全身操作研究
通过模仿学习实现全身移动操作对机器人在多样环境和复杂任务中的技能泛化至关重要。然而,这一目标面临重大挑战,尤其是在处理复杂观测、实现稳健泛化和生成连贯动作方面。为此,我们提出DSPv2,一种新型策略架构。DSPv2引入有效的编码方案,将3D空间特征与多视角2D语义特征对齐,该融合使策略在保持精细感知的同时实现广泛泛化。此外,我们将密集策略范式扩展至全身移动操作领域,验证其在全机器人平台上生成连贯且精确动作的有效性。大量实验表明,该方法在任务表现和泛化能力上均显著优于现有方法。
原文摘要 · Abstract (English)
Learning whole-body mobile manipulation via imitation is essential for generalizing robotic skills to diverse environments and complex tasks. However, this goal is hindered by significant challenges, particularly in effectively processing complex observation, achieving robust generalization, and generating coherent actions. To address these issues, we propose DSPv2, a novel policy architecture. DSPv2 introduces an effective encoding scheme that aligns 3D spatial features with multi-view 2D semantic features. This fusion enables the policy to achieve broad generalization while retaining the fine-grained perception necessary for precise control. Furthermore, we extend the Dense Policy paradigm to the whole-body mobile manipulation domain, demonstrating its effectiveness in generating coherent and precise actions for the whole-body robotic platform. Extensive experiments show that our method significantly outperforms existing approaches in both task performance and generalization ability. Project page is available at: https://selen-suyue.github.io/DSPv2Net/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。