arXiv:2506.09278cs.CVcs.LG2025-06NeurIPS被引 36

统一训练模型同时提升光流与大跨度图像匹配精度

UFM: A Simple Path towards Unified Dense Correspondence with Flow

  • 用通用Transformer直接回归像素位移向量
  • 比当前最优光流方法高28%准确率,速度快6.7倍
  • 首次证明统一模型在两类任务上均优于专用模型

密集图像对应是视觉里程计、三维重建、物体关联和重识别等应用的核心。历史上,宽基线场景与光流估计长期被分别处理,尽管目标相同。本文提出统一流与匹配模型(UFM),在源图与目标图中可见像素上进行统一训练。UFM采用简单通用的Transformer架构,直接回归(u,v)位移向量,相比以往基于代价体的粗到细方法更易训练且对大位移更准确。UFM比当前最优光流方法Unimatch高28%准确率,同时误差低62%,速度达密集宽基线匹配器RoMa的6.7倍。这是首个证明统一训练在两大领域均超越专用方法的研究,实现了快速通用的图像对应,为多模态、长程及实时对应任务开辟新方向。

原文摘要 · Abstract (English)

Dense image correspondence is central to many applications, such as visual odometry, 3D reconstruction, object association, and re-identification. Historically, dense correspondence has been tackled separately for wide-baseline scenarios and optical flow estimation, despite the common goal of matching content between two images. In this paper, we develop a Unified Flow & Matching model (UFM), which is trained on unified data for pixels that are co-visible in both source and target images. UFM uses a simple, generic transformer architecture that directly regresses the (u,v) flow. It is easier to train and more accurate for large flows compared to the typical coarse-to-fine cost volumes in prior work. UFM is 28% more accurate than state-of-the-art flow methods (Unimatch), while also having 62% less error and 6.7x faster than dense wide-baseline matchers (RoMa). UFM is the first to demonstrate that unified training can outperform specialized approaches across both domains. This result enables fast, general-purpose correspondence and opens new directions for multi-modal, long-range, and real-time correspondence tasks.

图像对应光流估计Transformer统一模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。