构建视频理解新基准,测试模型像人一样看懂复杂视频
Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding
- 用真实短视频+对抗性问题评估模型理解力
- 模型表现远低于人类,差距明显
- 适合研究视频推理与鲁棒性的团队使用
人类智能需要准确性和鲁棒性,前者是后者的基础。在视频理解中,准确性确保对视觉内容的正确解读,鲁棒性则保证在挑战性条件下表现稳定。尽管视频大语言模型(video LLMs)取得进展,现有基准未能充分反映这些模型与人类智能在视频解读中的准确性和鲁棒性差距。我们提出视频思维测试(Video-TT),用于评估视频LLMs是否能像人类一样有效理解现实世界视频。Video-TT包含1,000个YouTube Shorts视频,每个视频配一个开放问题和四个对抗性问题,以检验视觉与叙事复杂性。评估显示,视频LLMs与人类表现存在显著差距。
原文摘要 · Abstract (English)
Human intelligence requires correctness and robustness, with the former being foundational for the latter. In video understanding, correctness ensures the accurate interpretation of visual content, and robustness maintains consistent performance in challenging conditions. Despite advances in video large language models (video LLMs), existing benchmarks inadequately reflect the gap between these models and human intelligence in maintaining correctness and robustness in video interpretation. We introduce the Video Thinking Test (Video-TT), to assess if video LLMs can interpret real-world videos as effectively as humans. Video-TT reflects genuine gaps in understanding complex visual narratives, and evaluates robustness against natural adversarial questions. Video-TT comprises 1,000 YouTube Shorts videos, each with one open-ended question and four adversarial questions that probe visual and narrative complexity. Our evaluation shows a significant gap between video LLMs and human performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。