arXiv:2511.20162cs.CVcs.AI2025-11

测试视频大模型对物理交互起止点的判断能力,发现其依赖表面模式而非真实物理理解。

Action Without Interaction: Probing the Physical Foundations of Video LMMs via Contact-Release Detection

论文配图:Action Without Interaction: Probing the Physical Foundations of Video LMMs via Contact-Release Detection
图 1 · 摘自论文原文
  • 构建2万+视频交互事件数据集,标注接触与脱离时刻
  • 模型能准确识别动作和物体,但定位交互起止帧错误率超70%
  • 揭示模型存在‘捷径学习’,缺乏对物理接触的底层认知

大型多模态模型(LMMs)在图像和视频任务中表现日益出色,能够详细描述物体、环境及动态行为。本研究探究这些模型的语义理解是否基于真实视觉输入。针对手物交互视频序列,我们要求模型判断交互开始或结束的时间与位置。为此,我们引入首个大规模数据集,包含来自Something-Something-V2的20,000+条标注交互事件,由250名AMTurk标注员标注核心交互事件,特别是对象与主体何时发生接触(contact)或脱离(release)。我们测试了GPT、Gemini和Qwen等主流LMM,在每段仅含一个事件的短视频中定位这些事件。结果显示,尽管模型能可靠识别目标物体与动作,却普遍无法正确判断交互起止帧,且在场景中定位物理事件的能力较差。这种脱节表明,尽管模型擅长系统1式直觉识别(命名动作与物体),却缺乏系统2所需的物理基础推理能力,无法真正将动态场景建立在物理现实之上。

原文摘要 · Abstract (English)

Large multi-modal models (LMMs) show increasing performance in realistic visual tasks for images and, more recently, for videos. For example, given a video sequence, such models are able to describe in detail objects, the surroundings and dynamic actions. In this study, we explored the extent to which these models ground their semantic understanding in the actual visual input. Specifically, given sequences of hands interacting with objects, we asked models when and where the interaction begins or ends. For this purpose, we introduce a first of its kind, large-scale dataset with more than 20K annotated interactions on videos from the Something-Something-V2 dataset. 250 AMTurk human annotators labeled core interaction events, particularly when and where objects and agents become attached (`contact') or detached (`release'). We asked SoTA LMMs, including GPT, Gemini and Qwen to locate these events in short videos, each with a single event. The results show that while models reliably name target objects and identify actions, they exhibit a form of `shortcut learning' where semantic success masks a failure in physical grounding. Specifically, they consistently fail to identify the frame where the interaction begins or ends and poorly localize the physical event within the scene. This disconnect suggests that while LMMs excel at System 1 intuitive pattern recognition (naming the action and objects), they lack the System 2 cognitive foundations required to reason about physical primitives like `contact' and `release', hence truly ground dynamic scenes in physical reality.

视频理解物理推理大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。