arXiv:2603.03956cs.CVcs.AI2026-03被引 1

通过合成异模图像对,提升单应性估计在未知模态下的泛化能力。

Towards Generalized Multimodal Homography Estimation

  • 从单张图像生成带真实偏移的异模图像对用于训练。
  • 在跨模态测试中显著提升估计精度与鲁棒性。
  • 适合需要跨域泛化的视觉定位任务研究者。

监督与无监督的单应性估计方法依赖于特定模态的图像对以获得高精度,但在面对未见过的模态时性能显著下降。为解决此问题,我们提出一种训练数据合成方法,仅需一张输入图像即可生成具有真实偏移的非对齐图像对。该方法在保留结构信息的同时,引入多样化的纹理与颜色。这些合成数据使模型具备更强的鲁棒性和跨域泛化能力。此外,我们设计了一种网络,充分融合多尺度信息,并将颜色信息从特征表示中解耦,从而提升估计精度。大量实验表明,所提数据合成方法有效提升泛化性能,网络设计也得到验证。

原文摘要 · Abstract (English)

Supervised and unsupervised homography estimation methods depend on image pairs tailored to specific modalities to achieve high accuracy. However, their performance deteriorates substantially when applied to unseen modalities. To address this issue, we propose a training data synthesis method that generates unaligned image pairs with ground-truth offsets from a single input image. Our approach renders the image pairs with diverse textures and colors while preserving their structural information. These synthetic data empower the trained model to achieve greater robustness and improved generalization across various domains. Additionally, we design a network to fully leverage cross-scale information and decouple color information from feature representations, thus improving estimation accuracy. Extensive experiments show that our training data synthesis method improves generalization performance. The results also confirm the effectiveness of the proposed network.

单应性估计跨模态数据合成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。