提出真实场景下的视频目标分割新任务与数据集。
Show Me When and Where: Towards Referring Video Object Segmentation in the Wild
- 构建基于未剪辑视频的多场景数据集,要求定位目标出现的时间与位置。
- 新数据集时长是现有数据集的7倍,包含1120个真实视频。
- 提出OMFormer模型,在复杂场景下实现精准时空定位,适合实际应用研究。
指称视频目标分割(RVOS)在计算机视觉中因广泛应用而受到关注。现有设置通常使用精心剪辑的视频,且指称对象在所有帧中均出现,这无法反映真实挑战。该简化设定使方法只需预测目标位置,无需判断其出现时间。为此,本文引入面向真实场景的RVOS新设定,构建了基于YouTube未剪辑视频的新基准数据集YoURVOS,包含1,120个真实视频,时长比现有数据集多7倍,场景更丰富。该数据集要求模型不仅定位目标位置,还需判断其出现时间。为建立基线,提出物体级多模态变换器(OMFormer),通过编码物体级多模态交互,实现高效全局时空定位。实验表明,现有VOS方法在该数据集上表现显著下降,尤其当目标缺失帧增多时;而我们的OMFormer始终表现稳定。YoURVOS为推动实际应用中的RVOS方法发展提供了必要基准。
原文摘要 · Abstract (English)
Referring video object segmentation (RVOS) has recently generated great popularity in computer vision due to its widespread applications. Existing RVOS setting contains elaborately trimmed videos, with text-referred objects always appearing in all frames, which however fail to fully reflect the realistic challenges of this task. This simplified setting requires RVOS methods to only predict where objects, with no need to show when the objects appear. In this work, we introduce a new setting towards in-the-wild RVOS. To this end, we collect a new benchmark dataset using Youtube Untrimmed videos for RVOS - YoURVOS, which contains 1,120 in-the-wild videos with 7 times more duration and scenes than existing datasets. Our new benchmark challenges RVOS methods to show not only where but also when objects appear in videos. To set a baseline, we propose Object-level Multimodal TransFormers (OMFormer) to tackle the challenges, which are characterized by encoding object-level multimodal interactions for efficient and global spatial-temporal localisation. We demonstrate that previous VOS methods struggle on our YoURVOS benchmark, especially with the increase of target-absent frames, while our OMFormer consistently performs well. Our YoURVOS dataset offers an imperative benchmark, which will push forward the advancement of RVOS methods for practical applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。