arXiv:2510.02311cs.CVcs.LG2025-10被引 6

从视频中预测物体弹性、黏度等动态物理属性

Inferring Dynamic Physical Properties from Video Foundation Models

  • 用视频生成与自监督模型,通过视觉提示提取动态物理特征
  • 在合成与真实数据上验证,模型性能接近但略逊于理想基准
  • 适合对物理模拟、视频理解感兴趣的科研人员

我们研究从视频中预测动态物理属性的任务,重点关注需依赖时间信息的属性:弹跳物体的弹性、流动液体的黏度,以及滑动物体与表面间的动态摩擦。为此,我们做出三项贡献:(i) 构建了每种物理属性的新视频数据集,包含合成训练/测试集及真实世界评估集;(ii) 探索三种推断方法:(a) 基于经典计算机视觉的视觉线索(理想基准);(b) 在预训练视频生成与自监督模型上使用视觉提示和可训练提示向量进行交叉注意力读出;(c) 针对多模态大语言模型(MLLMs)的提示策略;(iii) 实验表明,基于生成式(DynamiCrafter)或自监督(V-JEPA-2)训练的视频基础模型表现相近,虽低于理想基准,而当前MLLMs表现较差,但可通过合适提示提升。数据集、模型与代码已公开。

原文摘要 · Abstract (English)

We study the task of predicting dynamic physical properties from videos. More specifically, we consider physical properties that require temporal information to be inferred: elasticity of a bouncing object, viscosity of a flowing liquid, and dynamic friction of an object sliding on a surface. To this end, we make the following contributions: (i) We collect a new video dataset for each physical property, consisting of synthetic training and testing splits, as well as a real split for real world evaluation. (ii) We explore three ways to infer the physical property from videos: (a) an oracle method where we supply the visual cues that intrinsically reflect the property using classical computer vision techniques; (b) a simple read out mechanism using a visual prompt and trainable prompt vector for cross-attention on pre-trained video generative and self-supervised models; and (c) prompt strategies for Multi-modal Large Language Models (MLLMs). (iii) We show that a video foundation model trained in a generative (DynamiCrafter) or trained in a self-supervised manner (V-JEPA-2) achieve a generally similar performance, though behind that of the oracle, and that MLLMs are currently inferior to the other models, though their performance can be improved through suitable prompting. The dataset, model, and code are available at https://www.robots.ox.ac.uk/~vgg/research/idpp/.

视频理解物理属性基础模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。