CVPR 2025挑战赛揭示复杂视频像素级理解新进展
PVUW 2025 Challenge Report: Advances in Pixel-level Understanding of Complex Videos in the Wild
- 设立两大赛道,分别聚焦复杂场景与运动引导的语言视频分割
- 推出更贴近真实场景的新数据集,推动算法在野外环境下的性能提升
- 适合关注视频理解、多模态分割与真实世界应用的研究者
本报告全面回顾了与CVPR 2025联合举办的第四届野外复杂视频像素级理解挑战赛(PVUW 2025)。挑战赛设两个赛道:MOSE专注于复杂场景下的视频对象分割,MeViS则针对运动引导、语言驱动的视频分割任务。两个赛道均引入全新且更具挑战性的数据集,以更真实地反映现实场景。通过详尽的评估与分析,报告揭示了当前复杂视频分割的最先进水平及新兴研究趋势。更多信息请访问研讨会官网:https://pvuw.github.io/。
原文摘要 · Abstract (English)
This report provides a comprehensive overview of the 4th Pixel-level Video Understanding in the Wild (PVUW) Challenge, held in conjunction with CVPR 2025. It summarizes the challenge outcomes, participating methodologies, and future research directions. The challenge features two tracks: MOSE, which focuses on complex scene video object segmentation, and MeViS, which targets motion-guided, language-based video segmentation. Both tracks introduce new, more challenging datasets designed to better reflect real-world scenarios. Through detailed evaluation and analysis, the challenge offers valuable insights into the current state-of-the-art and emerging trends in complex video segmentation. More information can be found on the workshop website: https://pvuw.github.io/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。