提出新框架,让模型看清被遮挡物体的完整形状。
A2VIS: Amodal-Aware Approach to Video Instance Segmentation
- 用时空先验掩码头融合可见与不可见部分信息
- 在VIS和MOT任务中提升遮挡场景下的追踪精度
- 适合需要精确识别遮挡物体的应用场景
遮挡仍是多对象跟踪(MOT)和视频实例分割(VIS)等视频实例级任务的主要挑战。本文提出一种新框架Amodal-Aware Video Instance Segmentation(A2VIS),引入非可视(amodal)表征以实现对视频中物体可见与被遮挡部分的可靠且全面的理解。核心思想是:通过时空维度感知非可视分割,可获得更稳定的对象信息流。当物体部分或完全被遮挡时,非可视分割相比可见分割在时间轴上变化更小、更一致。因此,可将所有片段中的非可视与可见信息整合为一个全局实例原型。为有效解决视频非可视分割问题,提出时空先验非可视掩码头,利用片段内可见信息并提取片段间非可视特征。大量实验与消融研究证明,A2VIS在识别与追踪具有完整形态理解的对象实例方面,在MOT和VIS任务中表现优异。
原文摘要 · Abstract (English)
Handling occlusion remains a significant challenge for video instance-level tasks like Multiple Object Tracking (MOT) and Video Instance Segmentation (VIS). In this paper, we propose a novel framework, Amodal-Aware Video Instance Segmentation (A2VIS), which incorporates amodal representations to achieve a reliable and comprehensive understanding of both visible and occluded parts of objects in a video. The key intuition is that awareness of amodal segmentation through spatiotemporal dimension enables a stable stream of object information. In scenarios where objects are partially or completely hidden from view, amodal segmentation offers more consistency and less dramatic changes along the temporal axis compared to visible segmentation. Hence, both amodal and visible information from all clips can be integrated into one global instance prototype. To effectively address the challenge of video amodal segmentation, we introduce the spatiotemporal-prior Amodal Mask Head, which leverages visible information intra clips while extracting amodal characteristics inter clips. Through extensive experiments and ablation studies, we show that A2VIS excels in both MOT and VIS tasks in identifying and tracking object instances with a keen understanding of their full shape.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。