让视频模型能自我进化,通过反事实干预验证未来行动是否可行。
Autonomous Video Generation with Counterfactual Controllability for Self-Evolving World Models

- 构建生成-绑定-验证-提炼闭环,实现自主视频生成
- 在无人机与机械臂上验证了风阻、延迟等约束下的有效性
- 强调生成视频的实用性而非画质,适合具身智能研究者
大规模视频生成模型被视作世界模型,因其可从视觉数据中学习丰富的时空规律。然而,理想的模型应具备自我演进的生成能力。仅追求视觉逼真度不足以判断想象的未来对特定具身代理是否物理可行,也难以提供环境反馈以推动改进。为此,本文提出自主视频生成框架,以反事实可控性为评估标准:即生成干预条件下的未来、绑定到具身约束、在分布外条件下验证,并将存活分支提炼为决策变量。提出四阶段闭环优化(生成、绑定、验证、提炼)及对应四项指标:新颖性、一致性、分布外(OOD)表现和效率。以无人机和机械臂为例,系统扰动风速、感知限制、执行延迟、接触动力学与恢复约束进行验证。核心观点是:自主视频生成不应仅看视频质量,而应评估其能否在反事实干预和多种具身约束下提升有效行动能力。
原文摘要 · Abstract (English)
Large-scale video generation models are increasingly described as world models because they can learn rich spatiotemporal regularities from visual data. However, we argue that an ideal world model should benefit in a self-evolving generative character. Traditional visually plausible predictions alone are not enough to establish whether an imagined future is physically actionable for a particular embodied agent, failing to provide informative feedback from environments for self-evolving improvement. To realize self-evolving world models, this article proposes the concept of autonomous video generation, which is evaluated through counterfactual controllability, i.e., the ability to i) generate intervention-conditioned futures, ii) bind these future frames to embodiment constraints, iii) verify them under distribution shifts, and iv) distil surviving branches into compact variables for decision-making. We formalize a four-stage closed-loop optimization of Generation, Binding, Verification and Distillation, together with four corresponding evaluation metrics: novelty, consistency, out-of-distribution (OOD) and efficiency. We further discuss two examples, i.e., drones and manipulators, as early embodied testbeds where wind, sensing limits, actuation delay, contact dynamics and recovery constraints can be systematically perturbed and verified. The central claim is that the framework of autonomous video generation for self-evolving world models should not be judged by video fidelity alone, but by whether the generated frames improve valid action under counterfactual interventions and various embodiment constraints.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。