arXiv:2511.13715cs.CV2025-11AAAI被引 3

解决视频跨镜头分割难题,提升模型在复杂切换下的泛化能力。

Segment Anything Across Shots: A Method and Benchmark

  • 设计过渡模仿数据增强策略,用单镜头数据学习跨镜头分割。
  • 提出SAAS模型,在多镜头视频上实现领先分割性能。
  • 构建新基准Cut-VOS,支持高频率转场的精细评估。

本文聚焦多镜头半监督视频对象分割(MVOS),旨在基于初始掩码对跨多个镜头的视频中的目标对象进行分割。现有方法主要针对单镜头视频,难以处理镜头间断导致的分割断裂,限制了实际应用。为此,我们提出一种过渡模仿数据增强策略(TMA),利用单镜头数据实现跨镜头泛化,缓解多镜头标注数据稀缺问题;并构建了可有效检测与理解镜头切换的跨镜头分割模型SAAS。为支持MVOS研究,我们引入新的基准Cut-VOS,包含密集掩码标注、多样物体类别和高频镜头切换。在YouMVOS和Cut-VOS上的大量实验表明,所提SAAS模型在复杂过渡场景下实现了当前最优性能。代码与数据集已公开于https://henghuiding.com/SAAS/。

原文摘要 · Abstract (English)

This work focuses on multi-shot semi-supervised video object segmentation (MVOS), which aims at segmenting the target object indicated by an initial mask throughout a video with multiple shots. The existing VOS methods mainly focus on single-shot videos and struggle with shot discontinuities, thereby limiting their real-world applicability. We propose a transition mimicking data augmentation strategy (TMA) which enables cross-shot generalization with single-shot data to alleviate the severe annotated multi-shot data sparsity, and the Segment Anything Across Shots (SAAS) model, which can detect and comprehend shot transitions effectively. To support evaluation and future study in MVOS, we introduce Cut-VOS, a new MVOS benchmark with dense mask annotations, diverse object categories, and high-frequency transitions. Extensive experiments on YouMVOS and Cut-VOS demonstrate that the proposed SAAS achieves state-of-the-art performance by effectively mimicking, understanding, and segmenting across complex transitions. The code and datasets are released at https://henghuiding.com/SAAS/.

视频分割跨镜头半监督数据增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。