用3D语义场提升扩散策略对未知物体的泛化能力
GenDP: 3D Semantic Fields for Category-Level Generalizable Diffusion Policy
- 通过多视角RGBD生成3D语义描述场,融合几何与语义信息
- 在未见物体上成功率从20%提升至93%,跨类别泛化效果显著
- 适合需要处理多样化物体布局的机器人操作任务研究者
基于扩散的策略在执行复杂机器人操作任务中表现优异,但缺乏对几何与语义的显式表征,常限制其在未见物体和布局上的泛化能力。为增强扩散策略的泛化性,本文提出新框架,通过3D语义场显式融入空间与语义信息。利用大型基础视觉模型从多视角RGBD观测生成3D描述场,并与参考描述场对比得到语义场。该方法显式建模几何与语义,使策略在需类别级泛化的任务中具备强泛化能力,解决几何歧义并关注细微几何特征。我们在涉及多类可动物体及不同形状纹理实例的8个任务上评估,结果表明,该方法将扩散策略在未见实例上的平均成功率从20%提升至93%。此外,我们提供详细分析与可视化,解释性能提升来源,并说明方法如何泛化到新实例。
原文摘要 · Abstract (English)
Diffusion-based policies have shown remarkable capability in executing complex robotic manipulation tasks but lack explicit characterization of geometry and semantics, which often limits their ability to generalize to unseen objects and layouts. To enhance the generalization capabilities of Diffusion Policy, we introduce a novel framework that incorporates explicit spatial and semantic information via 3D semantic fields. We generate 3D descriptor fields from multi-view RGBD observations with large foundational vision models, then compare these descriptor fields against reference descriptors to obtain semantic fields. The proposed method explicitly considers geometry and semantics, enabling strong generalization capabilities in tasks requiring category-level generalization, resolving geometric ambiguities, and attention to subtle geometric details. We evaluate our method across eight tasks involving articulated objects and instances with varying shapes and textures from multiple object categories. Our method demonstrates its effectiveness by increasing Diffusion Policy's average success rate on unseen instances from 20% to 93%. Additionally, we provide a detailed analysis and visualization to interpret the sources of performance gain and explain how our method can generalize to novel instances.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。