系统梳理视频场景解析进展,揭示关键挑战与未来方向
A Comprehensive Survey on Video Scene Parsing:Advances, Challenges, and Prospects
- 按架构演进梳理从传统特征到基础模型的五类任务方法
- 指出时间闪烁、遮挡导致身份切换等共性失败模式
- 适合关注视频理解系统设计与评估的研究者参考
视频场景解析(VSP)致力于密集视频理解,要求每帧中每个像素被分割、每个区域被命名,且对象身份在时间上保持一致。本文综述了视频语义分割(VSS)、视频实例分割(VIS)、视频全景分割(VPS)、视频跟踪与分割(VTS)及开放词汇视频分割(OVVS)五项任务的最新进展。将文献按架构演进组织为一条主线:从手工设计的运动与外观特征,到全卷积、注意力机制与查询驱动结构,再到近期的基础模型方法,并分析各类型如何建模时序上下文、保持身份一致性及平衡精度与效率。对比当前主流数据集、评价指标与基准趋势。除归纳方法外,重点揭示跨领域的设计权衡与重复出现的失效模式,包括时间闪烁、遮挡引起的身份切换、长尾类别问题以及标注-容量-延迟之间的张力。最后提出迈向鲁棒、高效、开放世界的VSP系统的关键研究方向。
原文摘要 · Abstract (English)
Video Scene Parsing (VSP) studies dense video understanding, where every pixel in each frame must be segmented, each region must be named, and each object identity must remain coherent over time. This survey reviews recent progress in VSP across five tasks, spanning Video Semantic Segmentation (VSS), Video Instance Segmentation (VIS), Video Panoptic Segmentation (VPS), Video Tracking \& Segmentation (VTS), and Open-Vocabulary Video Segmentation (OVVS). We organize the literature as one architectural arc, running from hand-crafted motion and appearance cues through fully convolutional, attention-based and query-based designs to recent foundation-model approaches, and we trace how each family models temporal context, preserves identity, and balances accuracy against efficiency. We then compare the datasets, metrics and benchmark trends that shape current evaluation. Beyond cataloguing methods, we foreground the design trade-offs and recurring failure modes that cut across the field, namely temporal flicker, occlusion-induced identity switches, long-tail categories, and the annotation--capacity--latency tension. We close with open directions towards robust, efficient and open-world VSP systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。