arXiv:2505.04864cs.CVcs.AI2025-05ICCV被引 2

通过自回归迭代优化,实现图像对齐在复杂情况下的高精度

Auto-regressive transformation for image alignment

  • 自回归管道逐级细化变换场,多尺度特征聚焦关键区域
  • 在平面图像上显著优于现有方法,3D场景也表现接近顶尖水平
  • 适合处理低纹理、大形变等困难场景,适用于高精度对齐任务

现有图像对齐方法在特征稀疏区域、极端尺度与视场差异、大形变情况下常表现不佳,导致精度不足。为提升鲁棒性,我们提出一种自回归变换(ART)方法,通过多尺度图像表示中迭代精炼变换场,并聚焦关键区域。网络在每一尺度上利用随机采样点优化变换参数,结合交叉注意力层引导,确保在特征匮乏条件下仍能准确对齐。大量实验表明,ART在平面图像上显著优于当前最优方法,在3D场景图像上也达到可比性能,展现出强大且通用的高精度图像对齐能力。

原文摘要 · Abstract (English)

Existing methods for image alignment struggle in cases involving feature-sparse regions, extreme scale and field-of-view differences, and large deformations, often resulting in suboptimal accuracy. Robustness to these challenges can be improved through iterative refinement of the transform field while focusing on critical regions in multi-scale image representations. We thus propose Auto-Regressive Transformation (ART), a novel method that iteratively estimates the coarse-to-fine transformations through an auto-regressive pipeline. Leveraging hierarchical multi-scale features, our network refines the transform field parameters using randomly sampled points at each scale. By incorporating guidance from the cross-attention layer, the model focuses on critical regions, ensuring accurate alignment even in challenging, feature-limited conditions. Extensive experiments demonstrate that ART significantly outperforms state-of-the-art methods on planar images and achieves comparable performance on 3D scene images, establishing it as a powerful and versatile solution for precise image alignment.

图像对齐自回归多尺度变换场

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。