构建游戏视频物理常识违例评测集,提升模型对物理常识的理解能力。
PhysGame: Uncovering Physical Commonsense Violations in Gameplay Videos
- 设计多维度物理违例游戏视频数据集,覆盖四大领域12类常识
- 发现开源视频模型性能远低于闭源模型,差距显著
- 提出增强型模型PhysVLM,支持物理知识注入与对抗训练
近期视频大模型在动态视觉内容理解方面取得进展,而游戏视频因其常含违背物理常识的漏洞,成为评估模型物理常识理解能力的理想基准。本文提出首个针对游戏视频中物理常识违例的评测基准PhysGame,包含880个视频,涵盖力学、运动学、光学和材料属性四个基础领域,以及12种具体物理常识。通过广泛评估多种先进视频大模型,发现当前开源模型性能明显落后于闭源模型。为弥补这一差距,我们构建了包含140,057个问答对的指令微调数据集PhysInstruct,并提出包含34,358组训练样本的偏好优化数据集PhysDPO,其中低质量响应分别基于误导性标题(元信息劫持)、少帧数(时间劫持)和低分辨率(空间劫持)生成。基于该数据集体系,我们提出物理知识增强型视频大模型PhysVLM。在PhysGame及通用视频理解基准上的实验表明,PhysVLM达到领先水平。
原文摘要 · Abstract (English)
Recent advancements in video-based large language models (Video LLMs) have witnessed the emergence of diverse capabilities to reason and interpret dynamic visual content. Among them, gameplay videos stand out as a distinctive data source, often containing glitches that defy physics commonsense. This characteristic renders them an effective benchmark for assessing the under-explored capability of physical commonsense understanding in video LLMs. In this paper, we propose PhysGame as a pioneering benchmark to evaluate physical commonsense violations in gameplay videos. PhysGame comprises 880 videos associated with glitches spanning four fundamental domains (i.e., mechanics, kinematics, optics, and material properties) and across 12 distinct physical commonsense. Through extensively evaluating various state-ofthe-art video LLMs, our findings reveal that the performance of current open-source video LLMs significantly lags behind that of proprietary counterparts. To bridge this gap, we curate an instruction tuning dataset PhysInstruct with 140,057 question-answering pairs to facilitate physical commonsense learning. In addition, we also propose a preference optimization dataset PhysDPO with 34,358 training pairs, where the dis-preferred responses are generated conditioned on misleading titles (i.e., meta information hacking), fewer frames (i.e., temporal hacking) and lower spatial resolutions (i.e., spatial hacking). Based on the suite of datasets, we propose PhysVLM as a physical knowledge-enhanced video LLM. Extensive experiments on both physical-oriented benchmark PhysGame and general video understanding benchmarks demonstrate the state-ofthe-art performance of PhysVLM.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。