用中间图像链模拟物理过程,让AI模型更懂时间演化规律。
Chain of Time: In-Context Physical Simulation with Image Generation Models
- 通过生成一系列中间图像实现时序物理推理
- 显著提升主流图像生成模型的物理模拟性能
- 适合研究视觉语言模型内在推理机制的学者
我们提出一种受认知启发的新方法——「Chain of Time」,用于提升和解释视觉语言模型中的物理模拟能力。该方法在推理阶段生成一系列中间图像,灵感来自机器学习中的上下文推理及人类的心理模拟。无需额外微调即可应用,涵盖二维图形模拟与真实三维视频等场景,测试了速度、加速度、流体动力学及动量守恒等多种物理特性。实验表明,采用该方法可显著提升当前最先进的图像生成模型性能。此外,我们分析了模型在各时间步的模拟状态,揭示了传统评估难以发现的动态机制,包括对速度、重力、碰撞等随时间演化的物理属性的建模能力。同时,也识别出模型虽能模拟物理过程,却难以从输入图像中推断特定物理参数的局限性。
原文摘要 · Abstract (English)
We propose a novel cognitively-inspired method to improve and interpret physical simulation in vision-language models. Our ``Chain of Time" method involves generating a series of intermediate images during a simulation, and it is motivated by in-context reasoning in machine learning, as well as mental simulation in humans. Chain of Time is used at inference time, and requires no additional fine-tuning. We apply the Chain-of-Time method to synthetic and real-world domains, including 2-D graphics simulations and natural 3-D videos. These domains test a variety of particular physical properties, including velocity, acceleration, fluid dynamics, and conservation of momentum. We found that using Chain-of-Time simulation substantially improves the performance of a state-of-the-art image generation model. Beyond examining performance, we also analyzed the specific states of the world simulated by an image model at each time step, which sheds light on the dynamics underlying these simulations. This analysis reveals insights that are hidden from traditional evaluations of physical reasoning, including cases where an image generation model is able to simulate physical properties that unfold over time, such as velocity, gravity, and collisions. Our analysis also highlights particular cases where the image generation model struggles to infer particular physical parameters from input images, despite being capable of simulating relevant physical processes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。