用几何与内容特征增强线查询,提升相机标定精度。
SOFI: Multi-Scale Deformable Transformer for Camera Calibration with Enhanced Line Queries
- 引入多尺度可变形注意力,融合线的几何与内容特征
- 在三个数据集上优于现有方法,推理速度保持高效
- 适合需要高精度相机参数的3D重建与图像合成任务
相机标定旨在估计如天顶消失点和地平线等相机参数,为3D渲染、虚拟现实效果和图像中物体插入提供支持。基于Transformer的模型虽表现良好,但缺乏跨尺度交互。本文提出SOFI(多尺度可变形变换器用于带增强线查询的相机标定),通过结合线的内容特征与几何特征改进了CTRL-C和MSCC中的线查询。SOFI使变换器模型能够采用多尺度可变形注意力机制,促进主干网络生成特征图间的跨尺度交互。在Google Street View、Horizon Line in the Wild和Holicity数据集上,SOFI均超越现有方法,同时保持具有竞争力的推理速度。
原文摘要 · Abstract (English)
Camera calibration consists of estimating camera parameters such as the zenith vanishing point and horizon line. Estimating the camera parameters allows other tasks like 3D rendering, artificial reality effects, and object insertion in an image. Transformer-based models have provided promising results; however, they lack cross-scale interaction. In this work, we introduce \textit{multi-Scale defOrmable transFormer for camera calibratIon with enhanced line queries}, SOFI. SOFI improves the line queries used in CTRL-C and MSCC by using both line content and line geometric features. Moreover, SOFI's line queries allow transformer models to adopt the multi-scale deformable attention mechanism to promote cross-scale interaction between the feature maps produced by the backbone. SOFI outperforms existing methods on the \textit {Google Street View}, \textit {Horizon Line in the Wild}, and \textit {Holicity} datasets while keeping a competitive inference speed.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。