CVPR 2026挑战赛聚焦像素级视频理解,探索多模态在复杂场景下的应用。
Report of the 5th PVUW Challenge: Towards More Diverse Modalities in Pixel-Level Understanding

- 设立三类赛道,涵盖密集遮挡追踪、语言驱动定位与音频驱动分割。
- 引入未公开的高难度数据集,评估模型在真实世界条件下的表现。
- 适合关注多模态视频理解与鲁棒场景分析的研究者参考。
本文总结了2026年在CVPR 2026举办的“野外像素级视频理解”(PVUW)挑战赛的目标、数据集及顶尖方法。该挑战赛旨在评估模型在高度非受限条件下的性能。2026年版本设立了三个专项赛道:针对密集遮挡和严重遮挡场景的MOSE赛道;基于运动相关语言表达进行目标定位的MeViS-Text赛道;以及首次推出的、以声音驱动为目标分割的MeViS-Audio赛道。通过引入此前未发布的高难度数据集,并分析参赛者提交的前沿多模态解决方案,本报告展示了社区在鲁棒视频场景理解方面的最新技术进展,并指明了未来发展方向。
原文摘要 · Abstract (English)
This report summarizes the objectives, datasets, and top-performing methodologies of the 2026 Pixel-level Video Understanding in the Wild (PVUW) Challenge, hosted at CVPR 2026, which evaluates state-of-the-art models under highly unconstrained conditions. To provide a comprehensive assessment, the 2026 edition features three specialized tracks: the MOSE track for tracking objects within densely cluttered and severely occluded scenarios; the MeViS-Text track for localizing targets via motion-focused linguistic expressions; and the newly inaugurated MeViS-Audio track, which pioneers acoustic-driven object segmentation. By introducing previously unreleased challenging data and analyzing the cutting-edge, multimodal solutions submitted by participants, this report highlights the community's latest technical advancements and charts promising future directions for robust video scene comprehension.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。