arXiv:2511.20886cs.CV2025-11被引 13

让SAM2跨视角匹配物体,用双提示生成器解决视差难题。

V$^{2}$-SAM: Marrying SAM2 with Multi-Prompt Experts for Cross-View Object Correspondence

  • 设计双提示生成器,分别处理几何与外观线索。
  • 在Ego-Exo4D等三个数据集上达最新最好性能。
  • 适合需要跨视角物体对齐的机器人与视频分析场景。

跨视角物体对应(如第一人称与第三人称视角对应)旨在建立不同视角下同一物体的一致关联,但因视角和外观差异大,现有分割模型(如SAM2)难以直接应用。为此,我们提出V2-SAM,一种统一的跨视角物体对应框架,通过两个互补的提示生成器将SAM2从单视角分割拓展至跨视角对应。其中,基于DINOv3特征的跨视角锚点提示生成器(V2-Anchor)建立几何感知对应,并首次实现对SAM2在跨视角场景下的坐标式提示;跨视角视觉提示生成器(V2-Visual)通过新颖的视觉提示匹配器,从特征与结构双重角度对齐第一人称与第三人称表征。为有效融合双提示优势,采用多专家设计,并引入后处理循环一致性选择器(PCCS),根据循环一致性自适应选择最可靠专家。大量实验验证了V2-SAM的有效性,在Ego-Exo4D(第一人称-第三人称物体对应)、DAVIS-2017(视频物体跟踪)和HANDAL-X(机器人可用的跨视角对应)上均达到新最优性能。

原文摘要 · Abstract (English)

Cross-view object correspondence, exemplified by the representative task of ego-exo object correspondence, aims to establish consistent associations of the same object across different viewpoints (e.g., egocentric and exocentric). This task poses significant challenges due to drastic viewpoint and appearance variations, making existing segmentation models, such as SAM2, difficult to apply directly. To address this, we present V2-SAM, a unified cross-view object correspondence framework that adapts SAM2 from single-view segmentation to cross-view correspondence through two complementary prompt generators. Specifically, the Cross-View Anchor Prompt Generator (V2-Anchor), built upon DINOv3 features, establishes geometry-aware correspondences and, for the first time, enables coordinate-based prompting for SAM2 in cross-view scenarios, while the Cross-View Visual Prompt Generator (V2-Visual) enhances appearance-guided cues via a novel visual prompt matcher that aligns ego-exo representations from both feature and structural perspectives. To effectively exploit the strengths of both prompts, we further adopt a multi-expert design and introduce a Post-hoc Cyclic Consistency Selector (PCCS) that adaptively selects the most reliable expert based on cyclic consistency. Extensive experiments validate the effectiveness of V2-SAM, achieving new state-of-the-art performance on Ego-Exo4D (ego-exo object correspondence), DAVIS-2017 (video object tracking), and HANDAL-X (robotic-ready cross-view correspondence).

跨视角匹配SAM2视觉对齐机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。