构建首个评估虚构视频生成与理解能力的基准,推动视频模型突破现实限制。
Impossible Videos
- 提出IPV-Bench基准,涵盖4大领域14类违背物理规律的视频场景。
- 发现当前视频生成模型难以准确遵循抽象指令创作不可能画面。
- 适合研究视频生成、认知推理与超越现实约束的AI系统开发者。
当前合成视频多用于补充真实世界数据的稀缺性与多样性,但主要局限于模拟现实场景,对不可能、反事实及违背现实的视频内容探索不足。本文旨在回答两个问题:1)现有视频生成模型能否有效遵循提示生成不可能内容?2)现有视频理解模型是否具备理解此类视频的能力?为此,我们提出了IPV-Bench,一个全新的基准,用于评估和推动视频生成与理解的发展。该基准基于全面的分类体系,包含4个领域、14个类别,覆盖违反物理、生物、地理或社会规律的多样化场景。基于此分类,构建了提示集以测试生成模型的指令遵循与创造力;同时,设计了视频评估集,用于检验视频-大语言模型在理解不可能视频时对时间动态与世界知识的推理能力。综合评估揭示了现有模型的局限,并为下一代视频模型指明了发展方向。
原文摘要 · Abstract (English)
Synthetic videos nowadays is widely used to complement data scarcity and diversity of real-world videos. Current synthetic datasets primarily replicate real-world scenarios, leaving impossible, counterfactual and anti-reality video concepts underexplored. This work aims to answer two questions: 1) Can today's video generation models effectively follow prompts to create impossible video content? 2) Are today's video understanding models good enough for understanding impossible videos? To this end, we introduce IPV-Bench, a novel benchmark designed to evaluate and foster progress in video understanding and generation. IPV-Bench is underpinned by a comprehensive taxonomy, encompassing 4 domains, 14 categories. It features diverse scenes that defy physical, biological, geographical, or social laws. Based on the taxonomy, a prompt suite is constructed to evaluate video generation models, challenging their prompt following and creativity capabilities. In addition, a video benchmark is curated to assess Video-LLMs on their ability of understanding impossible videos, which particularly requires reasoning on temporal dynamics and world knowledge. Comprehensive evaluations reveal limitations and insights for future directions of video models, paving the way for next-generation video models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。