arXiv:2504.11218cs.CV2025-04被引 14

用3D高斯点云提升物体功能区域识别精度,解决传统方法泛化差问题。

3DAffordSplat: Efficient Affordance Reasoning with 3D Gaussians

  • 构建首个基于3D高斯的多模态功能推理数据集,含2万多实例。
  • 提出新模型使功能识别准确率显著提升,跨场景泛化能力强。
  • 适合做机器人操作、具身智能等任务的研究者参考。

3D功能推理对将人类指令与3D物体的功能区域关联至关重要,有助于实现具身智能中的精准任务操作。然而,现有方法主要依赖稀疏点云,因对坐标变化敏感且数据本身稀疏,导致泛化性与鲁棒性不足。相比之下,3D高斯点阵(3DGS)通过密集连续表示实现高质量实时渲染,计算开销小,能更好捕捉精细功能细节。但其潜力受限于缺乏大规模专用数据集。为此,我们提出3DAffordSplat——首个面向3DGS的功能推理大规模多模态数据集,包含23,677个高斯实例、8,354个点云实例和6,631个手动标注的功能标签,涵盖21类物体与18种功能类型。基于此,我们设计了AffordSplatNet,一种专为3DGS表示优化的功能推理模型。该模型引入交叉模态结构对齐模块,利用结构一致性先验对齐点云与3DGS表示,显著提升识别准确率。大量实验表明,3DAffordSplat显著推动3DGS领域功能学习,AffordSplatNet在已见与未见场景中均优于现有方法,展现出强泛化能力。

原文摘要 · Abstract (English)

3D affordance reasoning is essential in associating human instructions with the functional regions of 3D objects, facilitating precise, task-oriented manipulations in embodied AI. However, current methods, which predominantly depend on sparse 3D point clouds, exhibit limited generalizability and robustness due to their sensitivity to coordinate variations and the inherent sparsity of the data. By contrast, 3D Gaussian Splatting (3DGS) delivers high-fidelity, real-time rendering with minimal computational overhead by representing scenes as dense, continuous distributions. This positions 3DGS as a highly effective approach for capturing fine-grained affordance details and improving recognition accuracy. Nevertheless, its full potential remains largely untapped due to the absence of large-scale, 3DGS-specific affordance datasets. To overcome these limitations, we present 3DAffordSplat, the first large-scale, multi-modal dataset tailored for 3DGS-based affordance reasoning. This dataset includes 23,677 Gaussian instances, 8,354 point cloud instances, and 6,631 manually annotated affordance labels, encompassing 21 object categories and 18 affordance types. Building upon this dataset, we introduce AffordSplatNet, a novel model specifically designed for affordance reasoning using 3DGS representations. AffordSplatNet features an innovative cross-modal structure alignment module that exploits structural consistency priors to align 3D point cloud and 3DGS representations, resulting in enhanced affordance recognition accuracy. Extensive experiments demonstrate that the 3DAffordSplat dataset significantly advances affordance learning within the 3DGS domain, while AffordSplatNet consistently outperforms existing methods across both seen and unseen settings, highlighting its robust generalization capabilities.

3D高斯功能推理具身智能多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。