arXiv:2605.26383cs.CV2026-05

用多阶段融合提升零样本厨房视频物体重识别效果

Zero-Shot Object Re-Identification in Egocentric Kitchen Videos via Multi-Stage SAM3 Feature Fusion

论文配图:Zero-Shot Object Re-Identification in Egocentric Kitchen Videos via Multi-Stage SAM3 Feature Fusion
图 1 · 摘自论文原文
  • 以SAM3为核心,分四阶段融合视觉与几何特征
  • 在EPIC-Kitchens上达52.8% mAP,比基线提升7.5%
  • 适合无标注数据的厨房场景物体追踪与识别

由于视角快速变化、频繁遮挡、场景杂乱及类内外观差异大,头戴式厨房视频中的物体重识别(ReID)极具挑战。物体可能离开再进入视野,且实例多样性高而标注有限,使监督式ReID难以扩展,推动了零样本方法的发展。本文在EPIC-Kitchens基准上研究零样本物体ReID,目标是仅使用预训练视觉特征匹配活跃的食品与厨具实例。我们评估了五种先进特征提取器:CLIP、DINOv2、DreamSim、I-JEPA和SAM3,发现零样本方法表现不佳,最佳基线仅达45.3% mAP。随后提出增强型SAM3 ReID流水线,一种基于SAM3分割的核心零样本多阶段方法。第一阶段用SAM3抑制背景杂波;第二阶段融合SAM3、DINOv2与CLIP的嵌入,生成归一化描述符;第三阶段结合掩码形状交并比(IoU)优化余弦相似度,保证几何一致性;第四阶段采用k-互惠重排序。完整流程将性能提升7.5% mAP,达到52.8%。

原文摘要 · Abstract (English)

Object re-identification (ReID) in egocentric kitchen videos is challenging due to rapid viewpoint changes, frequent occlusions, cluttered scenes, and large intra-class appearance variations. Objects may leave and re-enter the field of view, and the large diversity of instances with limited annotations makes supervised ReID difficult to scale, motivating zero-shot approaches. We study zero-shot object ReID on the EPIC-Kitchens benchmark, where the goal is to match active food and kitchen-tool instances across frames using only pre-trained visual features. We first evaluate five state-of-the-art feature extractors, including Vision-Language Models (VLMs) - CLIP, DINOv2, DreamSim, I-JEPA, and SAM3 - and show that zero-shot methods fail, with the best baseline achieving only 45.3% mAP. We then propose an Enhanced SAM3 ReID Pipeline, a zero-shot multi-stage method built around SAM3 segmentation as the core component. Stage 1 uses SAM3 to suppress background clutter. Stage 2 fuses embeddings from SAM3, DINOv2, and CLIP into a single L2-normalized descriptor. Stage 3 augments cosine similarity with mask-shape IoU for geometric consistency, and Stage 4 applies k-reciprocal re-ranking. The full pipeline improves performance by 7.5% mAP to 52.8%.

物体重识别零样本SAM3厨房视频

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。