构建首个视频理解领域泛化数据集,评测模型在真实场景下的鲁棒性。
VUDG: A Dataset for Video Understanding Domain Generalization
- 设计跨11个领域的视频数据集,覆盖三类分布偏移
- 9个主流模型在跨域测试中普遍性能下降,最差降幅超30%
- 适合研究模型鲁棒性、跨域迁移的学者与工程师
近年来,视频理解取得显著进展,主要得益于深度模型和大规模标注数据集的发展。然而,现有研究通常忽视真实视频应用中的固有领域偏移,导致视频理解领域的泛化能力研究不足。为此,我们提出视频理解领域泛化(VUDG)数据集,专用于评估视频理解中的领域泛化性能。VUDG包含来自11个不同领域的视频,涵盖三类领域偏移,并保持各领域间语义一致性,以确保评估的公平性与意义。我们采用多专家渐进式标注框架,为每段视频生成多项选择和开放式问答对。在9个代表性大型视频-语言模型(LVLMs)及若干传统视频问答方法上的广泛实验表明,大多数模型(包括顶尖的LVLMs)在领域偏移下均出现性能下降。这些结果凸显了VUDG带来的挑战,以及当前模型对数据分布偏移鲁棒性的差异。我们相信,VUDG将为未来视频理解领域泛化研究提供重要资源。
原文摘要 · Abstract (English)
Video understanding has made remarkable progress in recent years, largely driven by advances in deep models and the availability of large-scale annotated datasets. However, existing works typically ignore the inherent domain shifts encountered in real-world video applications, leaving domain generalization (DG) in video understanding underexplored. Hence, we propose Video Understanding Domain Generalization (VUDG), a novel dataset designed specifically for evaluating the DG performance in video understanding. VUDG contains videos from 11 distinct domains that cover three types of domain shifts, and maintains semantic similarity across different domains to ensure fair and meaningful evaluation. We propose a multi-expert progressive annotation framework to annotate each video with both multiple-choice and open-ended question-answer pairs. Extensive experiments on 9 representative large video-language models (LVLMs) and several traditional video question answering methods show that most models (including state-of-the-art LVLMs) suffer performance degradation under domain shifts. These results highlight the challenges posed by VUDG and the difference in the robustness of current models to data distribution shifts. We believe VUDG provides a valuable resource for prompting future research in domain generalization video understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。