用视觉语言模型构建可模拟的抽象世界,让虚拟环境更真实、可交互。
VDAWorld: World Modelling via VLM-Directed Abstraction and Simulation
- 用VLM自动选择视觉工具和物理引擎,生成可模拟的2D/3D场景
- 从静态场景推断潜在动态,预测未来状态,支持多种仿真任务
- 适合需要物理推理与交互控制的研究者,如机器人、游戏开发
生成式视频模型在世界建模中面临根本性挑战:常违背物理与逻辑规则,缺乏交互性,且作为黑箱模型难以构建结构化、可查询的世界。为此,我们提出一种新范式,将图像与文本对提炼为适于仿真的抽象表示。提出VDAWorld框架,其中视觉语言模型(VLM)作为智能代理,自主选择视觉工具构建有根基的2D或3D场景,并匹配相应的物理仿真器(如刚体、流体)。该框架可从静态场景推断潜在动力学,预测合理未来状态。实验表明,这种智能抽象与自适应仿真结合,使世界模型在多种场景下生成高质量仿真。我们在交互控制、反事实生成、物理与逻辑推理等应用中验证了其有效性,在多个基准上达到当前最优性能。
原文摘要 · Abstract (English)
Generative video models, a leading approach to world modelling, face fundamental limitations. They often violate physical and logical rules, lack interactivity, and operate as opaque black boxes ill-suited for building structured, queryable worlds. To overcome these challenges, we propose a new paradigm focused on distilling an image caption pair into a tractable, abstract representation optimized for simulation. We introduce VDAWorld, a framework where a Vision-Language Model (VLM) acts as an intelligent agent to orchestrate this process. The VLM autonomously constructs a grounded (2D or 3D) scene representation by selecting from a suite of vision tools, and accordingly chooses a compatible physics simulator (e.g., rigid body, fluid) to act upon it. VDAWorld can then infer latent dynamics from the static scene to predict plausible future states. Our experiments show that this combination of intelligent abstraction and adaptive simulation results in a versatile world model capable of producing high quality simulations across a wide range of scenarios. We demonstrate several applications of VDAWorld across interactive control and counterfactual generation, and physical and logical reasoning, achieving state-of-the-art results on several benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。