arXiv:2605.23861cs.LGcs.AI2026-05

用大模型实现零样本因果生成,可自动推理干预与反事实图像。

Leveraging Foundation Models for Causal Generative Modeling

论文配图:Leveraging Foundation Models for Causal Generative Modeling
图 1 · 摘自论文原文
  • 分三步构建因果生成流水线:概念提取、干预操纵、反事实生成。
  • 零样本完成因果发现与图像反事实生成,保持语义一致性。
  • 适合需要可解释生成的视觉推理任务,如医疗影像分析。

因果生成建模对于实现具备反事实推理能力的可靠透明AI系统至关重要。现有方法虽在生成模型训练中引入因果约束,但缺乏统一框架来利用预训练基础模型的零样本推理能力。本文提出FM-CGM,一种基于预训练基础模型的端到端视觉因果推理模块化框架。该框架通过三个核心组件实现因果流程:概念提取器、概念操纵器和反事实生成器。利用大型推理模型进行因果推断,结合文本到图像扩散模型完成生成,实现零样本因果发现、干预与反事实生成。进一步设计了因果语义引导(CSG)机制,基于交叉注意力确保语义干预传播至下游概念,同时保留不变区域。实验证明,该方法能识别合理因果结构,适用于忠实的反事实图像生成。

原文摘要 · Abstract (English)

Causal generative modeling is essential for developing reliable and transparent AI systems capable of counterfactual reasoning. While existing approaches focus on integrating causal constraints during the training of generative models, they often lack a unified framework to leverage the zero-shot reasoning capabilities of pretrained foundation models. We introduce FM-CGM, a modular framework for end-to-end visual causal reasoning using pretrained foundation models. FM-CGM formalizes the causal pipeline through three core components: a concept extractor, a concept manipulator, and a counterfactual generator. By leveraging a large reasoning model for causal inference and a text-to-image diffusion model for generation, our approach enables zero-shot causal discovery, intervention, and counterfactual generation. We then develop Causal Semantic Guidance (CSG), a cross-attention-based mechanism that ensures semantic interventions propagate to descendant concepts while preserving invariant regions. We empirically show that our approach can identify plausible causal structures and is suitable for faithful counterfactual image generation.

因果生成大模型反事实

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。