arXiv:2503.21820cs.CVeess.IV2025-03被引 3

统一预训练模型提升多模态图像特征匹配能力

UFM: Unified Feature Matching Pre-training with Multi-Modal Image Assistants

  • 设计可微调的多模态图像助手结构,适配跨模态匹配任务
  • 在多种图像模态上实现优异泛化性能,克服数据稀疏与不平衡问题
  • 适合需要跨模态对齐的视觉任务研究者使用

图像特征匹配是计算机视觉的基础任务,但在多模态应用中仍具挑战性,通常需针对特定数据集进行复杂训练。本文提出统一特征匹配预训练模型(UFM),旨在解决多种图像模态下的特征匹配难题。我们引入多模态图像助手(MIA)Transformer,一种可微调的结构,能有效处理多样化的特征匹配问题。UFM在同模态与跨模态匹配任务中均表现出强适应性。此外,我们设计了一种数据增强算法和分阶段预训练策略,以应对特定模态数据稀疏及模态分布不均的问题。实验表明,UFM在各类特征匹配任务中均展现出出色的泛化能力与性能。代码将发布于:https://github.com/LiaoYun0x0/UFM。

原文摘要 · Abstract (English)

Image feature matching, a foundational task in computer vision, remains challenging for multimodal image applications, often necessitating intricate training on specific datasets. In this paper, we introduce a Unified Feature Matching pre-trained model (UFM) designed to address feature matching challenges across a wide spectrum of modal images. We present Multimodal Image Assistant (MIA) transformers, finely tunable structures adept at handling diverse feature matching problems. UFM exhibits versatility in addressing both feature matching tasks within the same modal and those across different modals. Additionally, we propose a data augmentation algorithm and a staged pre-training strategy to effectively tackle challenges arising from sparse data in specific modals and imbalanced modal datasets. Experimental results demonstrate that UFM excels in generalization and performance across various feature matching tasks. The code will be released at:https://github.com/LiaoYun0x0/UFM.

特征匹配多模态预训练Transformer

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。