用因果提示引导扩散模型生成真实可信的视频假设场景。
Causally Steered Diffusion for Automated Video Counterfactual Generation
- 通过因果图嵌入提示词,引导扩散模型生成符合因果关系的视频反事实。
- 在真实人脸视频上实现最优因果有效性,同时保持时序一致性和视觉质量。
- 无需修改编辑系统,适用于任意黑箱视频编辑工具,适合媒体与医疗场景。
将文本到图像潜在扩散模型(LDM)应用于视频编辑虽具高视觉保真度和可控性,但难以维持视频数据生成过程中的因果关系。忽略因果依赖属性的修改可能导致不真实或误导性结果。本文提出一种因果忠实的反事实视频生成框架,将其建模为分布外(OOD)预测问题。通过将因果图中的先验因果知识编码进文本提示,并利用基于视觉-语言模型(VLM)的文本损失优化提示,驱动LDM潜空间捕捉形式上的反事实分布外变化,有效引导生成符合因果逻辑的替代方案。该框架命名为CSVC,对底层视频编辑系统无感,无需内部机制访问或微调。我们使用标准视频质量指标及因果有效性、最小性等反事实特异性评估标准进行验证。实验表明,CSVC通过提示驱动的因果引导,在LDM分布内生成了因果忠实的视频反事实,实现了当前最优的因果有效性,同时未牺牲时间一致性或视觉质量。由于兼容任何黑箱视频编辑系统,本框架在数字媒体、医疗等领域生成‘假如……会怎样’的逼真假设视频场景具有巨大潜力。
原文摘要 · Abstract (English)
Adapting text-to-image (T2I) latent diffusion models (LDMs) to video editing has shown strong visual fidelity and controllability, but challenges remain in maintaining causal relationships inherent to the video data generating process. Edits affecting causally dependent attributes often generate unrealistic or misleading outcomes if these relationships are ignored. In this work, we introduce a causally faithful framework for counterfactual video generation, formulated as an Out-of-Distribution (OOD) prediction problem. We embed prior causal knowledge by encoding the relationships specified in a causal graph into text prompts and guide the generation process by optimizing these prompts using a vision-language model (VLM)-based textual loss. This loss encourages the latent space of the LDMs to capture OOD variations in the form of counterfactuals, effectively steering generation toward causally meaningful alternatives. The proposed framework, dubbed CSVC, is agnostic to the underlying video editing system and does not require access to its internal mechanisms or fine-tuning. We evaluate our approach using standard video quality metrics and counterfactual-specific criteria, such as causal effectiveness and minimality. Experimental results show that CSVC generates causally faithful video counterfactuals within the LDM distribution via prompt-based causal steering, achieving state-of-the-art causal effectiveness without compromising temporal consistency or visual quality on real-world facial videos. Due to its compatibility with any black-box video editing system, our framework has significant potential to generate realistic 'what if' hypothetical video scenarios in diverse areas such as digital media and healthcare.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。