arXiv:2503.21765cs.CV2025-03综述被引 31

梳理视频生成中物理认知的演进,指出现有模型‘视觉真实但物理荒谬’的痛点。

Exploring the Evolution of Physics Cognition in Video Generation: A Survey

  • 从感知、被动认知到主动模拟,构建三层次物理认知框架
  • 指出当前模型违背物理规律,生成内容存在逻辑矛盾
  • 适合关注视频生成可控性与真实性的研究者参考

近期视频生成技术虽因扩散模型快速发展取得显著进展,但其在物理认知方面的缺陷日益凸显——生成内容常违反基本物理定律,陷入‘视觉真实但物理荒谬’的困境。研究者开始重视物理保真度的重要性,尝试将运动表征与物理知识等启发式认知融入生成系统,以模拟真实世界动态。针对该领域缺乏系统综述的现状,本文从认知科学视角梳理了物理认知在视频生成中的演化过程,提出三层次分类体系:1)生成的基础模式感知,2)生成中的被动物理知识认知,3)生成世界的主动认知与模拟,涵盖前沿方法、经典范式与基准测试。同时,强调该领域内在关键挑战,并勾勒未来研究路径,旨在推动学术界与产业界从‘视觉模仿’迈向‘类人物理理解’的新阶段。

原文摘要 · Abstract (English)

Recent advancements in video generation have witnessed significant progress, especially with the rapid advancement of diffusion models. Despite this, their deficiencies in physical cognition have gradually received widespread attention - generated content often violates the fundamental laws of physics, falling into the dilemma of ''visual realism but physical absurdity". Researchers began to increasingly recognize the importance of physical fidelity in video generation and attempted to integrate heuristic physical cognition such as motion representations and physical knowledge into generative systems to simulate real-world dynamic scenarios. Considering the lack of a systematic overview in this field, this survey aims to provide a comprehensive summary of architecture designs and their applications to fill this gap. Specifically, we discuss and organize the evolutionary process of physical cognition in video generation from a cognitive science perspective, while proposing a three-tier taxonomy: 1) basic schema perception for generation, 2) passive cognition of physical knowledge for generation, and 3) active cognition for world simulation, encompassing state-of-the-art methods, classical paradigms, and benchmarks. Subsequently, we emphasize the inherent key challenges in this domain and delineate potential pathways for future research, contributing to advancing the frontiers of discussion in both academia and industry. Through structured review and interdisciplinary analysis, this survey aims to provide directional guidance for developing interpretable, controllable, and physically consistent video generation paradigms, thereby propelling generative models from the stage of ''visual mimicry'' towards a new phase of ''human-like physical comprehension''.

视频生成物理认知扩散模型综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。