arXiv:2503.21581cs.CVcs.AI2025-03ICCV被引 2

用扩散模型联合估计相机参数与场景几何,提升真实环境下的校准精度。

AlignDiff: Learning Physically-Grounded Camera Alignment via Diffusion

  • 基于几何先验的扩散模型,同时优化相机内参、外参和畸变
  • 在真实数据集上将射线束角度误差降低约8.2度
  • 适合需要高精度3D感知的自动驾驶、机器人视觉场景

精确的相机标定是3D感知的基础任务,尤其在存在复杂光学畸变的真实环境中。现有方法通常依赖预矫正图像或标定板,限制了适用性和灵活性。本文提出AlignDiff框架,通过通用光线相机模型联合建模相机内参与外参。不同于以往侧重语义特征的方法,AlignDiff聚焦几何特征,更准确地建模局部畸变。我们设计了一种以几何先验为条件的扩散模型,实现相机畸变与场景几何的联合估计。为增强畸变预测能力,引入边缘感知注意力机制,聚焦图像边缘附近的几何特征而非语义内容。此外,为提升对真实拍摄图像的泛化能力,构建了一个包含超过三千个样本的光线追踪镜头数据库,覆盖多种镜头形态的固有畸变特性。实验表明,该方法在挑战性的真实世界数据集上显著降低了约8.2度的射线束角度误差,整体校准精度优于现有方法。

原文摘要 · Abstract (English)

Accurate camera calibration is a fundamental task for 3D perception, especially when dealing with real-world, in-the-wild environments where complex optical distortions are common. Existing methods often rely on pre-rectified images or calibration patterns, which limits their applicability and flexibility. In this work, we introduce a novel framework that addresses these challenges by jointly modeling camera intrinsic and extrinsic parameters using a generic ray camera model. Unlike previous approaches, AlignDiff shifts focus from semantic to geometric features, enabling more accurate modeling of local distortions. We propose AlignDiff, a diffusion model conditioned on geometric priors, enabling the simultaneous estimation of camera distortions and scene geometry. To enhance distortion prediction, we incorporate edge-aware attention, focusing the model on geometric features around image edges, rather than semantic content. Furthermore, to enhance generalizability to real-world captures, we incorporate a large database of ray-traced lenses containing over three thousand samples. This database characterizes the distortion inherent in a diverse variety of lens forms. Our experiments demonstrate that the proposed method significantly reduces the angular error of estimated ray bundles by ~8.2 degrees and overall calibration accuracy, outperforming existing approaches on challenging, real-world datasets.

相机标定扩散模型3D感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。