arXiv:2606.31077cs.CV2026-06中稿 · ECCV

用单视角图生成高质量多模态图像对,解决匹配数据稀缺问题

AnyMatch: Supercharging Universal Multi-Modal Image Matching with Large-Scale Single-View Images

  • 利用单目深度与3D重投影合成多视角多模态图像对
  • 在多个基准上微调后,匹配性能显著提升,泛化能力更强
  • 适合需要高精度几何一致性数据的研究者

多模态图像匹配对视觉定位和多传感器融合至关重要,但受限于带精确几何标注的大规模训练数据稀缺。现有真实数据集成本高、场景多样性不足,且存在SfM-MVS流程中的误差;而合成方法难以保持3D几何一致性或实现逼真外观。为此,我们提出AnyMatch框架,利用大量易获取的单视角图像,在极低成本下生成丰富的多模态训练数据。AnyMatch结合单目深度估计、3D重投影、基于扩散的修补和跨模态图像转换,合成具有3D几何一致性的多视图、多模态图像对。关键在于通过显式3D重投影提供严格符合几何一致性的标注,避免了SfM-MVS的误差累积。此外,AnyMatch具备强可扩展性,可通过调整输入和相机参数控制场景多样性和标注难度。我们构建了Any-syn——一个大规模合成多模态数据集。实验表明,使用Any-syn微调的匹配网络(如LoFTR、EDM、RoMa)在多模态基准上取得显著性能提升,表现出更优的泛化性和鲁棒性。

原文摘要 · Abstract (English)

Multi-modal image matching is essential for visual localization and multi-sensor fusion, but it is hindered by the scarcity of large-scale training data with precise geometric annotations. Existing real-world datasets suffer from prohibitive costs, limited scene diversity, and errors in SfM-MVS pipelines, while synthetic methods struggle to maintain 3D geometric consistency or achieve photorealistic appearance. To address this, we propose AnyMatch, a novel framework that leverages abundant, easily accessible single-view images at minimal cost to generate rich multi-modal training data. AnyMatch integrates monocular depth estimation, 3D reprojection, diffusion-based inpainting, and crossmodal image translation to synthesize multi-view, multi-modal image pairs with 3D geometric fidelity. Crucially, our method provides annotations that strictly adhere to 3D geometric consistency through explicit 3D reprojection, avoiding SfM-MVS error accumulation. Furthermore, AnyMatch offers strong scalability, enabling controllable scene diversity and annotation difficulty via adjustable input and camera parameters. We construct Any-syn, a large-scale synthetic multi-modal dataset using AnyMatch. Experimental results show that matching networks (e.g., LoFTR, EDM, RoMa) fine-tuned on Any-syn achieve substantial performance gains on multi-modal benchmarks, exhibiting superior generalization and robustness compared to models trained on existing data.

多模态匹配数据合成3D一致性图像生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。