arXiv:2606.18586cs.CVcs.AI2026-06

用原子物理变化解释视频中物体运动的因果过程。

APT: Atomic Physical Transitions for Causal Video-Language Understanding

论文配图:APT: Atomic Physical Transitions for Causal Video-Language Understanding
图 1 · 摘自论文原文
  • 提出原子物理转换(APTs)捕捉物体运动中的关键因果状态变化。
  • 在27,303个实例上训练,模型零样本召回率仅14%,说明现有模型忽略细节过程。
  • 新方法APT-Tune用少量参数让模型既懂物理过程又不丢事件理解能力。

物理事件的理解不仅依赖名称,更依赖其背后的因果状态变化。一个片段标签如“弹跳”虽正确,却掩盖了支撑力丧失、接触开始、反弹和稳定等过程。为此,我们引入原子物理转换(APTs):最小、时间局部化的状态变化,将可见线索与主动物理机制及前后动力学状态关联。APTs链将视频表示为有序的因果转换序列,取代单一事件标签:标签说明发生了什么,而APTs链解释为何发生。为使APTs可被视觉语言模型(VLMs)学习,我们构建了混合源数据,结合人工标注与模拟器真值,涵盖接触、重力、摩擦和旋转/稳定性共14种转换类型,包含27,303个带时间戳实例,覆盖1,246次试验。实验发现,当前VLMs在转换级物理理解上表现极差,零样本召回率最高仅14%,错误主要源于遗漏转换。直接微调虽提升检测,但导致事件级遗忘,表明模型只学会特定回答格式而非可复用的物理表征。因此我们提出APT-Tune:一种参数高效方法,通过图像垫感知监督、格式条件联合训练和机制条件域到类型解码,实现对格式鲁棒且物理根基牢固的APTs学习。仅在Qwen3-VL-2B上使用1100万LoRA参数,即可显著提升APTs召回率,并同时改善事件级视频迁移性能。结果表明,APTs不仅是新答案格式,更是对齐人类认知的因果监督信号。

原文摘要 · Abstract (English)

Physical events are not understood by their names alone, but by the causal state changes that compose them. A clip-level label such as "bounce" can be correct while hiding the process that makes the event physically valid, from support loss and contact onset to rebound and settling. To make this hidden process explicit, we introduce Atomic Physical Transitions (APTs): minimal, temporally localized state changes that bind a visible cue to an active physical mechanism and before/after dynamical regimes. An APT chain represents a video as an ordered causal transition sequence rather than a single aggregate event label: event labels tell what happened; APT chains explain why it happened. To make APTs learnable by VLMs, we construct mixed-source APT data from human annotations and simulator ground truth, covering 14 transition types across contact, gravity, friction, and rotation/stability, with 27,303 timed instances over 1,246 trials. Using this data, we find that current VLMs miss transition-level physics, with zero-shot recall at most 14% and errors dominated by missed transitions. Direct fine-tuning on APT chains improves transition detection but causes event-level forgetting, indicating that the model learns a specialized answer format rather than a reusable physical representation. We therefore propose APT-Tune, a parameter-efficient recipe that teaches VLMs to use causal transitions without forgetting how to answer video questions. It combines image-pad-aware supervision, format-conditional co-training, and mechanism-conditioned domain-to-type decoding to make APT learning format-robust and physically grounded. With only 11 M LoRA parameters on Qwen3-VL-2B, APT-Tune substantially improves APT recall while also improving event-level video transfer. These results show that APTs are not a new answer format, but a human-aligned causal supervision signal for physical video understanding.

视频理解因果推理物理建模多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。