arXiv:2506.06026cs.CV2025-06ICCV被引 9

跨视角物体掩码匹配新方法,显著提升多视角分割精度。

O-MaMa: Learning Object Mask Matching between Egocentric and Exocentric Views

  • 将跨视图分割转为掩码匹配任务,用编码器和注意力融合双视角特征。
  • 在Ego-Exo4D基准上,跨视角分割准确率提升22%至76%。
  • 仅用1%参数量即达顶级性能,适合资源受限场景使用。

理解多视角世界对协同智能系统至关重要,但跨视角共同物体分割仍是开放难题。本文提出O-MaMa,将跨图像分割重新定义为掩码匹配任务。方法包含:(1) 掩码-上下文编码器,从FastSAM候选掩码中提取DINOv2语义特征,获得判别性物体表征;(2) 自主-他主交叉注意力,融合多视角观测;(3) 掩码匹配对比损失,使跨视角特征在共享隐空间对齐;(4) 困难负邻域挖掘策略,增强模型对邻近物体的区分能力。在Ego-Exo4D对应性基准上,相对于官方基线,Ego2Exo与Exo2Ego IoU分别提升+22%和+76%;相较于现有最优方法,参数量仅为1%却仍取得+13%和+6%的提升。

原文摘要 · Abstract (English)

Understanding the world from multiple perspectives is essential for intelligent systems operating together, where segmenting common objects across different views remains an open problem. We introduce a new approach that re-defines cross-image segmentation by treating it as a mask matching task. Our method consists of: (1) A Mask-Context Encoder that pools dense DINOv2 semantic features to obtain discriminative object-level representations from FastSAM mask candidates, (2) an Ego$\leftrightarrow$Exo Cross-Attention that fuses multi-perspective observations, (3) a Mask Matching contrastive loss that aligns cross-view features in a shared latent space, and (4) a Hard Negative Adjacent Mining strategy to encourage the model to better differentiate between nearby objects. O-MaMa achieves the state of the art in the Ego-Exo4D Correspondences benchmark, obtaining relative gains of +22% and +76% in the Ego2Exo and Exo2Ego IoU against the official challenge baselines, and a +13% and +6% compared with the SOTA with 1% of the training parameters.

跨视角分割掩码匹配多模态小样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。