arXiv:2507.08400cs.CVcs.AI2025-07被引 1

用同一模型解决多种图像匹配任务,无需重新训练。

PanMatch: Unleashing the Potential of Large Vision Models for Unified Matching Models

  • 统一用位移估计框架处理多类匹配任务。
  • 在跨任务评测中超越UniMatch和Flow-Anything,零样本下表现优异。
  • 适合需要通用匹配能力的研究者与开发者。

本文提出PanMatch,一种适用于鲁棒对应匹配的通用基础模型。不同于以往依赖特定任务架构或领域微调的方法,我们发现任意两帧匹配任务均可通过统一的二维位移估计框架实现,且仅需相同模型权重。该方法避免了设计专用统一架构或任务特化集成模型,通过赋予位移估计算法前所未有的泛化能力实现多任务融合。为此,我们强调跨领域适用的鲁棒特征提取器的重要性,并提出特征转换流程,利用大视觉模型的通用特征,使匹配基线具备零样本跨视角匹配能力。我们构建了一个包含近180万样本的跨域数据集,涵盖立体匹配、光流与特征匹配任务,用于预训练PanMatch。实验表明,该模型在广泛领域和下游任务中使用相同权重即可表现出色,在跨任务评估中优于UniMatch和Flow-Anything,且在任务导向基准上达到多数先进方法水平。此外,PanMatch在雨天、卫星图像等异常场景中展现出前所未有的零样本性能,而现有稳健算法在此类场景中常失效。

原文摘要 · Abstract (English)

This work presents PanMatch, a versatile foundation model for robust correspondence matching. Unlike previous methods that rely on task-specific architectures and domain-specific fine-tuning to support tasks like stereo matching, optical flow or feature matching, our key insight is that any two-frame correspondence matching task can be addressed within a 2D displacement estimation framework using the same model weights. Such a formulation eliminates the need for designing specialized unified architectures or task-specific ensemble models. Instead, it achieves multi-task integration by endowing displacement estimation algorithms with unprecedented generalization capabilities. To this end, we highlight the importance of a robust feature extractor applicable across multiple domains and tasks, and propose the feature transformation pipeline that leverage all-purpose features from Large Vision Models to endow matching baselines with zero-shot cross-view matching capabilities. Furthermore, we assemble a cross-domain dataset with near 1.8 million samples from stereo matching, optical flow, and feature matching domains to pretrain PanMatch. We demonstrate the versatility of PanMatch across a wide range of domains and downstream tasks using the same model weights. Our model outperforms UniMatch and Flow-Anything on cross-task evaluations, and achieves comparable performance to most state-of-the-art task-specific algorithms on task-oriented benchmarks. Additionally, PanMatch presents unprecedented zero-shot performance in abnormal scenarios, such as rainy day and satellite imagery, where most existing robust algorithms fail to yield meaningful results.

图像匹配通用模型零样本大视觉模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。