arXiv:2608.07256cs.CV2026-08

用图像生成模型的语义先验,实现无需训练的3D物体方向标准化。

CANIS: Generation-Assisted 3D Canonicalization via an Image-Semantic Bridge

论文配图:CANIS: Generation-Assisted 3D Canonicalization via an Image-Semantic Bridge
图 1 · 摘自论文原文
  • 通过生成代理并利用图像-3D语义桥对齐特征。
  • 在任意旋转下提升分类、分割与对应任务性能。
  • 不依赖类别模板,适用于真实世界不完整扫描数据。

3D物体方向标准化是3D理解的基础。现有方法多依赖几何线索,但最终需语义有意义的方向。为此,我们提出CANIS——一种类别无关、生成辅助的框架,将冻结的图像到3D生成模型的语义方向先验引入3D标准化过程,无需特定训练或类别模板。CANIS首先从候选视角渲染输入物体,选取信息量大的视角,生成一个标准朝向的代理。生成过程中,由输入编码的稀疏结构潜在变量引导代理保持物体几何特征。随后,以选定图像作为输入与代理之间的语义桥梁:图像块识别代理上的语义区域,深度反投影定位输入上的对应区域。由此产生的语义锚点约束几何匹配,进而估计出使输入标准化的刚性变换。在合成基准上的实验验证了CANIS及其关键组件的有效性;在部分观测和OmniObject3D上的定性结果表明其适用于不完整和真实扫描数据。此外,CANIS在任意旋转下提升了下游的3D分类、部件分割和密集对应性能。

原文摘要 · Abstract (English)

Canonicalizing 3D object orientation is fundamental to 3D understanding and analysis. Existing approaches often rely on geometric cues, although 3D canonicalization ultimately requires a semantically meaningful orientation. To address this gap, we propose CANIS, a category-agnostic, generation-assisted framework that introduces the semantic orientation prior of a frozen image-to-3D generative model into 3D canonicalization, without canonicalization-specific training or category-specific templates. Specifically, CANIS first renders the input object from candidate viewpoints, selects an informative view, and generates a proxy in a canonical orientation. During generation, a sparse structural latent encoded from the input guides the proxy to preserve the geometry of an object. CANIS then uses the selected image as a semantic bridge between the input and the proxy. Image patches identify semantic regions on the proxy, and depth back-projection locates the corresponding regions on the input. The resulting semantic anchors constrain geometric matching, from which we estimate the rigid transformation that canonicalizes the input. Experiments on synthetic benchmarks validate CANIS and its key components, while qualitative results on partial observations and OmniObject3D suggest its applicability to incomplete and real-world scans. CANIS also improves downstream 3D classification, part segmentation, and dense correspondence under arbitrary rotations. Project page: https://kenkenzaii.github.io/Canis.

3D生成语义对齐姿态标准化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。