用游戏漏洞视频训练模型理解物理规律,效果优于传统方法。
Order from Chaos: Physical World Understanding from Glitchy Gameplay Videos
- 利用游戏中的物理错误作为训练信号,生成高质量问答数据。
- 在真实世界和通用场景下,模型推理能力提升1.9%至3.7%。
- 适合研究多模态物理理解、具身智能与数据增强的学者使用。
理解物理世界,包括物体运动、材料属性和因果交互,仍是人工智能的核心挑战。尽管多模态大语言模型(MLLM)展现出强大的通用推理能力,但在物理原理理解上仍远未达到人类水平。现有物理推理数据集要么依赖真实世界视频(标注成本高),要么基于合成模拟(现实感和多样性不足)。本文提出新范式:利用游戏视频中的视觉异常(即违反预设物理法则的漏洞)作为丰富的可扩展监督信号。我们构建了PhysGame数据集,包含140,057个以漏洞为中心的问答对,覆盖五个物理领域和十六个细粒度类别。通过结合游戏元信息(如标题、描述)设计提示策略,确保问答质量。同时,我们创建了GameBench基准,由880个专家标注的含漏洞游戏视频组成,用于评估物理推理能力。大量实验表明,经PhysGame微调的模型在真实世界推理任务中,使Qwen2.5VL在PhysBench上提升2.5%,在MVBench上提升1.9%;在GameBench上绝对提升达3.7%,显示其在识别物理不合理现象方面的更强鲁棒性。结果表明,从游戏异常中学习为提升多模态智能的物理理解提供了高效且可扩展的路径。
原文摘要 · Abstract (English)
Understanding the physical world, including object dynamics, material properties, and causal interactions, remains a core challenge in artificial intelligence. Although recent multi-modal large language models (MLLMs) have demonstrated impressive general reasoning capabilities, they still fall short of achieving human-level understanding of physical principles. Existing datasets for physical reasoning either rely on real-world videos, which incur high annotation costs, or on synthetic simulations, which suffer from limited realism and diversity. In this paper, we propose a novel paradigm that leverages glitches in gameplay videos, referring to visual anomalies that violate predefined physical laws, as a rich and scalable supervision source for physical world understanding. We introduce PhysGame, an meta information guided instruction-tuning dataset containing 140,057 glitch-centric question-answer pairs across five physical domains and sixteen fine-grained categories. To ensure data accuracy, we design a prompting strategy that utilizes gameplay metadata such as titles and descriptions to guide high-quality QA generation. Complementing PhysGame, we construct GameBench, an expert-annotated benchmark with 880 glitch-identified gameplay videos designed to evaluate physical reasoning capabilities. Extensive experiments show that PhysGame significantly enhances both Game2Real transferability, improving the real world physical reasoning performance of Qwen2.5VL by 2.5% on PhysBench, and Game2General transferability, yielding a 1.9% gain on the MVBench benchmark. Moreover, PhysGame-tuned models achieve a 3.7% absolute improvement on GameBench, demonstrating enhanced robustness in detecting physical implausibilities. These results indicate that learning from gameplay anomalies offers a scalable and effective pathway toward advancing physical world understanding in multimodal intelligence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。