arXiv:2412.09329cs.MMcs.AI2024-12被引 11

让视频语义分割模型能识别未见过的物体类别。

Towards Open-Vocabulary Video Semantic Segmentation

  • 融合时空信息,利用连续帧间关系提升理解。
  • 在VSPW和Cityscapes上实现对新类别的零样本分割。
  • 适合需要泛化到未知类别的视频分析场景。

视频语义分割是近期研究的重点,但现有模型在面对未知类别时表现不佳。为此,我们提出开放词汇视频语义分割(OV-VSS)任务,旨在准确分割广泛开放词汇类别中的每个像素,包括全新或未探索的类别。为提升性能,我们构建了基准方法OV2VSS,引入时空融合模块,利用连续帧间的时序关系;同时加入随机帧增强模块,扩展模型对整个视频语义上下文的理解;还结合视频文本编码,增强模型对视频中文字信息的解析能力。在VSPW和Cityscapes等基准数据集上的全面评估表明,该方法具备出色的零样本泛化能力,尤其在处理新类别时表现优异。结果验证了OV2VSS的有效性,在多个视频数据集上均实现了更优的分割性能。

原文摘要 · Abstract (English)

Semantic segmentation in videos has been a focal point of recent research. However, existing models encounter challenges when faced with unfamiliar categories. To address this, we introduce the Open Vocabulary Video Semantic Segmentation (OV-VSS) task, designed to accurately segment every pixel across a wide range of open-vocabulary categories, including those that are novel or previously unexplored. To enhance OV-VSS performance, we propose a robust baseline, OV2VSS, which integrates a spatial-temporal fusion module, allowing the model to utilize temporal relationships across consecutive frames. Additionally, we incorporate a random frame enhancement module, broadening the model's understanding of semantic context throughout the entire video sequence. Our approach also includes video text encoding, which strengthens the model's capability to interpret textual information within the video context. Comprehensive evaluations on benchmark datasets such as VSPW and Cityscapes highlight OV-VSS's zero-shot generalization capabilities, especially in handling novel categories. The results validate OV2VSS's effectiveness, demonstrating improved performance in semantic segmentation tasks across diverse video datasets.

视频分割开放词汇零样本时空建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。