arXiv:2605.27938cs.CV2026-05

让3D模型在不同实例间保持语义一致,提升重建准确性

SEMAGIC: Learning Semantically Consistent Deformable 3D Representations from In-the-Wild Images

论文配图:SEMAGIC: Learning Semantically Consistent Deformable 3D Representations from In-the-Wild Images
图 1 · 摘自论文原文
  • 用可变形模型发现类别级对应关系,以模板网格+形变场构建统一表示
  • 在SPair-71k数据集上语义对应准确率提升14.7%([email protected]
  • 适合需要稳定部件对应的任务,如3D形状分析、跨实例比较

从单视角野外图像中学习可变形3D物体模型已实现无需监督的出色3D形状重建。然而,这些模型是否捕捉了下游任务所需的语义结构仍不明确。我们发现,现有可变形重建方法尽管生成视觉上合理的几何体,但在不同实例间对应关系不稳定,且在语义对应基准测试中表现较差。本文提出SEMAGIC框架,从单视图野外图像中学习语义一致的可变形3D表示。不同于将重建视为最终目标,SEMAGIC将可变形建模作为发现类别级对应关系的机制。每个类别由一个规范模板网格和学习到的形变场表示,类似于自编码器,从图像特征重构实例几何,使顶点在不同实例间保持一致的语义含义。通过(i)特征级别一致性损失对齐规范与形变网格间的语义特征,以及(ii)基于顶点索引的形变,确保跨实例的语义对应关系。通过显式耦合几何形变与语义对齐,SEMAGIC生成的表示在类内变化下保持稳定的部件对应关系。实验表明,SEMAGIC在SPair-71k数据集上将可变形模型的语义对应准确率提升14.7%([email protected]),确立可变形模型作为有效的语义3D表示。

原文摘要 · Abstract (English)

Learning deformable 3D object models from single-view in-the-wild images has enabled impressive 3D shape reconstruction without supervision. However, it remains unclear whether these models capture the semantic structure required for downstream tasks. We find that existing deformable reconstruction approaches, despite producing visually plausible geometry, yield unstable correspondences across instances and perform poorly on semantic correspondence benchmarks. We introduce SEMAGIC, a framework for learning semantically consistent deformable 3D representations from single-view in-the-wild images. Rather than treating reconstruction as the end goal, SEMAGIC uses deformable modeling as a mechanism to discover category-level correspondences. Each category is represented by a canonical template mesh and a learned deformation field, functioning similarly to an autoencoder that reconstructs instance geometry from image features, enabling vertices to maintain consistent semantic meaning across instances. Semantic consistency is enforced during training through (i) a feature-level consistency loss aligning semantic features between canonical and deformed meshes, and (ii) vertex-index-conditioned deformation that preserves semantic correspondence across instances. By explicitly coupling geometric deformation with semantic alignment, SEMAGIC produces representations that maintain stable part correspondences across intra-category variation. Experiments demonstrate that SEMAGIC improves semantic correspondence of deformable models by +14.7 [email protected] on SPair-71k, establishing deformable models as effective semantic 3D representations.

3D重建语义对齐可变形模型形状分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。