arXiv:2510.08316cs.CV2025-10被引 3

用2D视觉模型的语义知识提升3D功能分割精度

Unlocking 3D Affordance Segmentation with 2D Semantic Knowledge

  • 通过跨模态对齐,让3D模型学习2D视觉模型的语义结构
  • 在多个数据集上准确率显著超越现有最佳方法
  • 适合做3D功能理解与人机交互的研究者参考

功能分割旨在将3D物体分解为具有不同功能的角色部分,使模型能推理物体间交互而非仅识别。现有方法多沿用3D语义分割或提示驱动框架,在几何线索弱或模糊时表现不佳,因稀疏点云提供的功能信息有限。为此,我们利用大规模2D视觉基础模型(VFMs)中蕴含的丰富语义知识,通过跨模态对齐机制引导3D表征学习。提出跨模态亲和力迁移(CMAT)预训练策略,强制3D编码器与提升后的2D特征所诱导的语义结构对齐。CMAT基于核心亲和力对齐目标,并辅以几何重建和特征多样性两个辅助损失,共同促进结构化且区分性强的特征学习。基于CMAT预训练的主干网络,采用轻量级功能分割器,通过高效交叉注意力接口注入文本或视觉提示,实现密集且提示感知的功能预测,同时保留预训练阶段建立的语义组织。大量实验表明,在准确率与效率上均持续优于先前最先进方法。

原文摘要 · Abstract (English)

Affordance segmentation aims to decompose 3D objects into parts that serve distinct functional roles, enabling models to reason about object interactions rather than mere recognition. Existing methods, mostly following the paradigm of 3D semantic segmentation or prompt-based frameworks, struggle when geometric cues are weak or ambiguous, as sparse point clouds provide limited functional information. To overcome this limitation, we leverage the rich semantic knowledge embedded in large-scale 2D Vision Foundation Models (VFMs) to guide 3D representation learning through a cross-modal alignment mechanism. Specifically, we propose Cross-Modal Affinity Transfer (CMAT), a pretraining strategy that compels the 3D encoder to align with the semantic structures induced by lifted 2D features. CMAT is driven by a core affinity alignment objective, supported by two auxiliary losses, geometric reconstruction and feature diversity, which together encourage structured and discriminative feature learning. Built upon the CMAT-pretrained backbone, we employ a lightweight affordance segmentor that injects text or visual prompts into the learned 3D space through an efficient cross-attention interface, enabling dense and prompt-aware affordance prediction while preserving the semantic organization established during pretraining. Extensive experiments demonstrate consistent improvements over previous state-of-the-art methods in both accuracy and efficiency.

3D分割功能理解跨模态学习视觉基础模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。