arXiv:2412.09511cs.CV2024-12CVPR被引 23

用2D大模型提升3D操作性识别的泛化与抗噪能力

GEAL: Generalizable 3D Affordance Learning with Cross-Modal Consistency

  • 融合2D预训练模型与3D点云,通过双分支架构实现跨模态对齐
  • 在多个数据集上优于现有方法,尤其在噪声和新类别下表现更优
  • 适合机器人操作、人机交互等需鲁棒感知的场景

从语义线索中识别3D物体的操作区域对机器人和人机交互至关重要。然而,现有3D操作性学习方法因标注数据有限且依赖仅关注几何编码的3D骨干网络,难以泛化且对真实世界噪声和数据损坏缺乏鲁棒性。我们提出GEAL框架,通过利用大规模预训练2D模型增强3D操作性学习的泛化与鲁棒性。采用双分支结构结合高斯点阵,建立3D点云与2D表示间的一致映射,实现从稀疏点云生成逼真2D渲染。引入粒度自适应融合模块和2D-3D一致性对齐模块,强化跨模态对齐与知识迁移,使3D分支可受益于2D模型丰富的语义与泛化能力。为全面评估鲁棒性,我们引入两个基于噪声的基准:PIAD-C和LASO-C。在公开数据集及自建基准上的大量实验表明,GEAL在已见与新类别物体以及受损数据上均持续优于现有方法,展现出多样条件下稳定可靠的预测能力。代码与噪声数据集已公开。

原文摘要 · Abstract (English)

Identifying affordance regions on 3D objects from semantic cues is essential for robotics and human-machine interaction. However, existing 3D affordance learning methods struggle with generalization and robustness due to limited annotated data and a reliance on 3D backbones focused on geometric encoding, which often lack resilience to real-world noise and data corruption. We propose GEAL, a novel framework designed to enhance the generalization and robustness of 3D affordance learning by leveraging large-scale pre-trained 2D models. We employ a dual-branch architecture with Gaussian splatting to establish consistent mappings between 3D point clouds and 2D representations, enabling realistic 2D renderings from sparse point clouds. A granularity-adaptive fusion module and a 2D-3D consistency alignment module further strengthen cross-modal alignment and knowledge transfer, allowing the 3D branch to benefit from the rich semantics and generalization capacity of 2D models. To holistically assess the robustness, we introduce two new corruption-based benchmarks: PIAD-C and LASO-C. Extensive experiments on public datasets and our benchmarks show that GEAL consistently outperforms existing methods across seen and novel object categories, as well as corrupted data, demonstrating robust and adaptable affordance prediction under diverse conditions. Code and corruption datasets have been made publicly available.

3D感知跨模态机器人鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。