无需训练即可精准匹配任意3D形状,抗噪能力强。
Articulating then Matching: Zero-Shot Shape Matching for Uncurated Data

- 用多视角渲染+几何一致性重建参数化形状模型
- 在非等距基准上误差降低73%~37%,最高处理20万顶点数据
- 适用于网格、点云、3D高斯等多种格式,测试时优化
在真实场景中寻找3D形状间的密集对应关系是基础但尚未解决的挑战。这些场景存在缺乏训练时间与样本、极端高分辨率未校准数据含拓扑扭曲、需处理多样3D表示等问题。本文提出ATM零样本框架,通过‘表述-匹配’范式一次性应对上述问题。不依赖内在几何特性,而是利用预训练视觉基础模型与参数化形状先验,从多视角渲染中估计参数化形状模型,并通过多视角几何一致性系统性校准。将不同输入映射至共享规范参数空间,自然建立鲁棒粗略对应,再经谱精修获得精确稠密映射。纯基于测试时优化的参数重建运行,无需对应训练数据,天然免疫连接性伪影,无缝兼容多种3D模态(网格、点云、3D高斯)。大量实验表明,方法在非等距基准上表现优异(平均测地误差:2.4-TOPKIDS,3.8-SMAL),较基线URSSM降低73%和37%;在高达20万顶点的野外原始扫描上展现出前所未有的鲁棒性,计算时间近似恒定且精度持续领先。
原文摘要 · Abstract (English)
Finding dense correspondences between 3D shapes is a fundamental yet unresolved challenge, especially in real-world environments. These environments present severe challenges, including the lack of time and sufficient samples for training, the prevalence of uncurated extreme-high resolution data with topological distortions, and the need to handle diverse 3D representations. In this paper, we present ATM, a zero-shot framework that requires no correspondence-specific training and robustly addresses these issues at once through an articulate-then-match paradigm. Rather than relying on intrinsic geometric properties, we leverage powerful pretrained vision foundation models and parametric shape priors to estimate parametric shape models from multi-view renderings, and systematically ground these estimations via multi-view geometric consistency. By mapping diverse inputs into a shared canonical parametric space, we inherently establish robust coarse correspondences that bypass topological noise, which are then refined into precise dense mappings via spectral refinement. Operating purely on test-time optimized parametric reconstructions, ATM requires no correspondence training data, is naturally immune to connectivity artifacts, and seamlessly handles diverse 3D modalities, including meshes, point clouds, and 3D Gaussians. Extensive experiments demonstrate that our method achieves strong results on non-isometric benchmarks (average geodesic errors of 2.4-TOPKIDS, 3.8-SMAL), reducing errors by 73% and 37% respectively compared to the baseline URSSM. Furthermore, it exhibits unprecedented robustness on in-the-wild raw scans of up to 200k vertices per shape while maintaining near-constant computation time and consistent superior accuracy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。