MOSEv2挑战真实复杂场景下的视频目标分割,推动算法落地。
MOSEv2: A More Challenging Dataset for Video Object Segmentation in Complex Scenes
- 构建5024段视频、70万+高质量标注,涵盖200类复杂场景
- 引入遮挡、小目标、恶劣天气等10余项真实难题,性能普遍下降超25%
- 适合研究真实世界视频理解的开发者和评测基准需求者
视频对象分割(VOS)旨在对视频中指定目标进行全程分割。尽管现有方法在DAVIS和YouTube-VOS等数据集上已达到90%以上交并比与平均精度(J&F),但这些数据集多包含显著、孤立且主导性物体,限制了其在真实场景中的泛化能力。为此,我们提出更复杂的MOSEv2数据集,作为MOSEv1的升级版,以推动真实复杂场景下的VOS研究。MOSEv2包含5,024个视频、701,976个高质量掩码,覆盖10,074个对象及200个类别。相比前代,该数据集显著提升场景复杂度,包括更频繁的物体消失/重现、严重遮挡与拥挤、小尺寸物体,以及新挑战如雨雪雾天气、低光照环境(夜间、水下)、多镜头序列、伪装物体、非物理目标(阴影、反射)及需外部知识的场景。我们在5种设置下测试20种代表性VOS方法,发现性能普遍下降;例如,SAM2从MOSEv1的76.4%降至50.9%。此外,9种视频目标跟踪方法也呈现类似下降趋势,表明该数据集对多任务构成挑战。分析表明,当前方法在真实复杂条件下仍存在明显不足。基于此,我们提出若干实用改进策略,有效提升模型表现。MOSEv2已公开发布于https://MOSE.video。
原文摘要 · Abstract (English)
Video object segmentation (VOS) aims to segment specified target objects throughout a video. Although state-of-the-art methods have achieved impressive performance (e.g., 90+% J&F) on benchmarks such as DAVIS and YouTube-VOS, these datasets primarily contain salient, dominant, and isolated objects, limiting their generalization to real-world scenarios. To bridge this gap, the coMplex video Object SEgmentation (MOSEv1) dataset was introduced to facilitate VOS research in complex scenes. Building on the foundations and insights of MOSEv1, we present MOSEv2, a significantly more challenging dataset designed to further advance VOS methods under real-world conditions. MOSEv2 consists of 5,024 videos and 701,976 high-quality masks for 10,074 objects across 200 categories. Compared to its predecessor, MOSEv2 introduces much greater scene complexity, including {more frequent object disappearance and reappearance, severe occlusions and crowding, smaller objects, as well as a range of new challenges such as adverse weather (e.g., rain, snow, fog), low-light scenes (e.g., nighttime, underwater), multi-shot sequences, camouflaged objects, non-physical targets (e.g., shadows, reflections), and scenarios requiring external knowledge.} We benchmark 20 representative VOS methods under 5 different settings and observe consistent performance drops on MOSEv2. For example, SAM2 drops from 76.4% on MOSEv1 to only 50.9% on MOSEv2. We further evaluate 9 video object tracking methods and observe similar declines, demonstrating that MOSEv2 poses challenges across tasks. These results highlight that despite strong performance on existing datasets, current VOS methods still fall short under real-world complexities. Based on our analysis of the observed challenges, we further propose several practical tricks that enhance model performance. MOSEv2 is publicly available at https://MOSE.video.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。