让视频生成模型能回答‘如果物体被移除会怎样’这类假设问题。
Counterfactual World Models via Digital Twin-conditioned Video Diffusion
- 用数字孪生结构化表示场景,实现对物体属性的精准干预。
- 在两个基准上达到当前最佳效果,验证了方法的有效性。
- 适合需要模拟多种假设场景的AI行为评估任务。
世界模型通过控制信号预测视觉观测的时间演化,使智能体能够通过前向仿真推理环境。然而,现有模型仅基于真实观测进行预测,难以回答如‘若移除某物体会发生什么’等反事实问题。本文提出反事实世界模型框架CWMDT,将标准视频扩散模型转化为可处理假设干预的模型。首先,构建场景的数字孪生,以结构化文本显式编码物体及其关系;其次,利用大语言模型推理干预如何随时间传播并改变场景;最后,以修改后的表示条件化视频扩散模型,生成反事实视觉序列。在两个基准上的评估表明,该方法性能达当前最优,证明数字孪生等替代视频表示可作为前向仿真世界模型的强大控制信号。
原文摘要 · Abstract (English)
World models learn to predict the temporal evolution of visual observations given a control signal, potentially enabling agents to reason about environments through forward simulation. Because of the focus on forward simulation, current world models generate predictions based on factual observations. For many emerging applications, such as comprehensive evaluations of physical AI behavior under varying conditions, the ability of world models to answer counterfactual queries, such as "what would happen if this object was removed?", is of increasing importance. We formalize counterfactual world models that additionally take interventions as explicit inputs, predicting temporal sequences under hypothetical modifications to observed scene properties. Traditional world models operate directly on entangled pixel-space representations where object properties and relationships cannot be selectively modified. This modeling choice prevents targeted interventions on specific scene properties. We introduce CWMDT, a framework to overcome those limitations, turning standard video diffusion models into effective counterfactual world models. First, CWMDT constructs digital twins of observed scenes to explicitly encode objects and their relationships, represented as structured text. Second, CWMDT applies large language models to reason over these representations and predict how a counterfactual intervention propagates through time to alter the observed scene. Third, CWMDT conditions a video diffusion model with the modified representation to generate counterfactual visual sequences. Evaluations on two benchmarks show that the CWMDT approach achieves state-of-the-art performance, suggesting that alternative representations of videos, such as the digital twins considered here, offer powerful control signals for video forward simulation-based world models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。