arXiv:2603.07652cs.CV2026-03

用视觉语言模型和图结构提升3D形状间语义对应精度

GLASS: Graph and Vision-Language Assisted Semantic Shape Correspondence

  • 融合视觉语言模型与图结构,实现无监督的3D形状语义对应
  • 在跨类别和非等距变形下误差降低37%以上,性能领先
  • 适合需要高精度3D匹配的机器人、动画生成等场景

建立3D形状间的密集对应关系对纹理转移、形状插值和机器人操作等下游任务至关重要。然而,在缺乏人工标注的情况下,尤其是在严重非等距变形和跨类别场景中,几何线索模糊时,学习这些映射仍具挑战性。传统函数映射方法因依赖等距性而表现受限。为此,本文提出GLASS框架,将几何谱分析与视觉-语言基础模型的丰富语义先验相结合。GLASS引入三项关键创新:(i) 视图一致性策略,从强大视觉基础模型中提取鲁棒的多视角视觉特征;(ii) 通过零样本3D分割将语言嵌入注入顶点描述符,捕捉高层部件语义;(iii) 基于测地线与拓扑关系的图辅助对比损失,强制区域间结构一致性(如源形状的‘头’ ↔ 目标形状的‘头’)。该设计使GLASS能在无真值监督下学习全局一致且语义连贯的映射。大量实验表明,GLASS在所有场景下均达到当前最优性能,在标准近等距任务上保持高精度,并显著提升挑战性场景下的表现。具体而言,在跨类别基准SNIS、非等距基准SMAL和TOPKIDS上,平均测地误差分别为0.21、4.5和5.6,相较URSSM基线分别降低57%、25%和37%。

原文摘要 · Abstract (English)

Establishing dense correspondence across 3D shapes is crucial for fundamental downstream tasks, including texture transfer, shape interpolation, and robotic manipulation. However, learning these mappings without manual supervision remains a formidable challenge, particularly under severe non-isometric deformations and in inter-class settings where geometric cues are ambiguous. Conventional functional map methods, while elegant, typically struggle in these regimes due to their reliance on isometry. To address this, we present GLASS, a framework that bridges the gap by integrating geometric spectral analysis with rich semantic priors from vision-language foundation models. GLASS introduces three key innovations: (i) a view-consistent strategy that enables robust multi-view visual feature extraction from powerful vision foundation models; (ii) the injection of language embeddings into vertex descriptors via zero-shot 3D segmentation, capturing high-level part semantics; and (iii) a graph-assisted contrastive loss that enforces structural consistency between regions (e.g., source's head'' $\leftrightarrow$ target's head'') by leveraging geodesic and topological relationships between regions. This design allows GLASS to learn globally coherent and semantically consistent maps without ground-truth supervision. Extensive experiments demonstrate that GLASS achieves state-of-the-art performance across all regimes, maintaining high accuracy on standard near-isometric tasks while significantly advancing performance in challenging settings. Specifically, it achieves average geodesic errors of 0.21, 4.5, and 5.6 on the inter-class benchmark SNIS and non-isometric benchmarks SMAL and TOPKIDS, reducing errors from URSSM baselines of 0.49, 6.0, and 8.9 by 57%, 25%, and 37%, respectively.

3D对应视觉语言模型无监督学习图神经网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。