arXiv:2412.05725cs.CVcs.AI2024-12CVPR被引 16

测试视觉语言模型在突发异常事件中的推理能力,发现其表现远逊于人类。

Black Swan: Abductive and Defeasible Video Reasoning in Unpredictable Events

  • 通过设计罕见事件场景,考察模型的反事实与可推翻推理能力。
  • 顶尖模型在关键任务上落后人类32%,暴露深层认知缺陷。
  • 适合研究AI认知局限、具身智能与鲁棒性评估的研究者。

视觉语言模型(VLMs)在常识推理,尤其是反事实推理和可推翻推理方面的能力仍不清晰。现有基准多聚焦典型视觉场景,难以区分模型表现是源于敏锐感知与推理,还是依赖统计记忆。我们主张聚焦视频中的非常规事件,以更清晰揭示VLM的核心能力。解释和理解此类分布外事件需模型超越基础模式识别与知识复述。为此,我们提出BlackSwanSuite,一个用于评估VLM在异常事件中进行反事实与可推翻推理的基准。任务通过限制视觉信息或引入新信息来挑战模型对隐藏事件的推断能力。该基准包含超过3,800道选择题、4,900道生成题和6,700道是非题,覆盖1,655个视频。我们对GPT-4o、Gemini 1.5 Pro及LLaVA-Video等先进VLM进行了广泛评估,发现其在这些任务上与人类存在最高达32%的性能差距。结果揭示当前VLM的关键局限,强调需改进模型架构与训练策略。数据与排行榜详见blackswan.cs.ubc.ca。

原文摘要 · Abstract (English)

The commonsense reasoning capabilities of vision-language models (VLMs), especially in abductive reasoning and defeasible reasoning, remain poorly understood. Most benchmarks focus on typical visual scenarios, making it difficult to discern whether model performance stems from keen perception and reasoning skills, or reliance on pure statistical recall. We argue that by focusing on atypical events in videos, clearer insights can be gained on the core capabilities of VLMs. Explaining and understanding such out-of-distribution events requires models to extend beyond basic pattern recognition and regurgitation of their prior knowledge. To this end, we introduce BlackSwanSuite, a benchmark for evaluating VLMs' ability to reason about unexpected events through abductive and defeasible tasks. Our tasks artificially limit the amount of visual information provided to models while questioning them about hidden unexpected events, or provide new visual information that could change an existing hypothesis about the event. We curate a comprehensive benchmark suite comprising over 3,800 MCQ, 4,900 generative and 6,700 yes/no questions, spanning 1,655 videos. After extensively evaluating various state-of-the-art VLMs, including GPT-4o and Gemini 1.5 Pro, as well as open-source VLMs such as LLaVA-Video, we find significant performance gaps of up to 32% from humans on these tasks. Our findings reveal key limitations in current VLMs, emphasizing the need for enhanced model architectures and training strategies. Our data and leaderboard is available at blackswan.cs.ubc.ca.

视频推理常识推理异常检测VLM评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。