从真实建筑图像中检测3D对称结构,解决单目视觉的定位难题。
ArchSym: Detecting 3D-Grounded Architectural Symmetries in the Wild

- 利用多视角匹配自动构建大规模建筑对称数据集ArchSym
- 提出基于符号距离图的3D对称定位方法,精确还原对称平面位置
- 首次在真实场景下实现高精度3D反射对称检测,适合建筑视觉任务
对称性检测是计算机视觉的基础问题,且可作为下游任务的强大先验。然而,现有基于学习的方法主要在物体中心或合成数据集上训练与评估,难以泛化至真实场景。此外,由于单目输入存在固有的尺度模糊性,导致3D平面定位为病态问题,多数方法仅预测平面方向。本文提出首个从单张真实场景RGB图像中检测3D接地反射对称性的框架,聚焦于建筑地标。我们引入两项关键创新:(1) 构建可扩展的数据标注流水线,通过利用多视角图像匹配,从SfM重建中自动提取大规模建筑对称数据集ArchSym;(2) 基于该数据集,设计单视图对称检测器,通过相对于预测场景几何定义的符号距离图参数化对称性,实现3D空间中对称性的精准定位。我们通过几何基准验证了标注流程的有效性,并证明所提检测器在新基准上显著优于现有最先进方法。
原文摘要 · Abstract (English)
Symmetry detection is a fundamental problem in computer vision, and symmetries serve as powerful priors for downstream tasks. However, existing learning-based methods for detecting 3D symmetries from single images have been almost exclusively trained and evaluated on object-centric or synthetic datasets, and thus fail to generalize to real-world scenes. Furthermore, due to the inherent scale ambiguity of monocular inputs, which makes localizing the 3D plane an ill-posed problem, many existing works only predict the plane's orientation. In this paper, we address these limitations by presenting the first framework for detecting 3D-grounded reflectional symmetries from single, in-the-wild RGB images, focusing on architectural landmarks. We introduce two key innovations: (1) a scalable data annotation pipeline to automatically curate a large-scale dataset of architectural symmetries, ArchSym, from SfM reconstructions by leveraging cross-view image matching; and building on the dataset, (2) a single-view symmetry detector that accurately localizes symmetries in 3D by parameterizing them as signed distance maps defined relative to predicted scene geometry. We validate our symmetry annotation pipeline against geometry-based alternatives and demonstrate that our symmetry detector significantly outperforms state-of-the-art baselines on our new benchmark.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。