arXiv:2512.13290cs.CVcs.AI2025-12

让扩散模型更懂物理规律,能应对新指令和复杂因果。

LINA: Learning INterventions Adaptively for Physical Alignment and Generalization in Diffusion Models

  • 通过因果图与探测数据集诊断模型缺陷。
  • 在图像视频生成中实现物理对齐与跨分布指令遵循。
  • 适合关注可控生成与因果推理的研究者。

扩散模型(DMs)在图像与视频生成上取得显著进展,但仍面临(1)物理对齐问题和(2)分布外(OOD)指令遵循困难。我们认为这些问题源于模型未能学习因果方向以及对因果因子的解耦,从而难以进行新组合。为此,我们提出因果场景图(CSG)和物理对齐探测集(PAP),用于诊断干预。分析得出三大关键发现:第一,扩散模型在提示中未明确指定元素时,难以进行多跳推理;第二,提示嵌入中包含可解耦的纹理与物理表征;第三,视觉因果结构主要在初始、计算受限的去噪步骤中建立。基于此,我们提出LINA(Learning INterventions Adaptively),一种新型框架,可自适应预测特定提示的干预策略,包含(1)提示与视觉隐空间中的定向引导,(2)重新分配的、基于因果意识的去噪调度。该方法在图像与视频扩散模型中同时实现物理对齐与分布外指令遵循,在挑战性因果生成任务及Winoground数据集上达到领先性能。

原文摘要 · Abstract (English)

Diffusion models (DMs) have achieved remarkable success in image and video generation. However, they still struggle with (1) physical alignment and (2) out-of-distribution (OOD) instruction following. We argue that these issues stem from the models' failure to learn causal directions and to disentangle causal factors for novel recombination. We introduce the Causal Scene Graph (CSG) and the Physical Alignment Probe (PAP) dataset to enable diagnostic interventions. This analysis yields three key insights. First, DMs struggle with multi-hop reasoning for elements not explicitly determined in the prompt. Second, the prompt embedding contains disentangled representations for texture and physics. Third, visual causal structure is disproportionately established during the initial, computationally limited denoising steps. Based on these findings, we introduce LINA (Learning INterventions Adaptively), a novel framework that learns to predict prompt-specific interventions, which employs (1) targeted guidance in the prompt and visual latent spaces, and (2) a reallocated, causality-aware denoising schedule. Our approach enforces both physical alignment and OOD instruction following in image and video DMs, achieving state-of-the-art performance on challenging causal generation tasks and the Winoground dataset. Our project page is at https://opencausalab.github.io/LINA.

扩散模型因果推理可控生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。