用因果最小性让生成模型的隐变量可解释且可控。
Beyond the Black Box: Identifiable Interpretation and Control in Generative Models via Causal Minimality
- 基于因果最小性原则,构建可解释的层级生成模型框架。
- 在文生图扩散模型中提取出内在的层次化概念图。
- 为模型精细调控提供可追踪的因果控制杠杆,适合研究者与开发者。
深度生成模型虽在图像、文本生成等领域带来革新,但普遍作为难以理解的“黑箱”,阻碍了人类对模型的理解、控制与对齐。尽管稀疏自编码器(SAEs)表现出显著的实证效果,却常缺乏理论保障,存在主观解读风险。本文旨在建立可解释生成模型的理论基础。我们证明,因果最小性原则——偏好最简因果解释——能使现代生成模型的隐变量具备明确的因果解释力和鲁棒的组件级可识别控制。我们提出一种新的分层选择模型理论框架,其中高层概念由低层变量受约束组合而成,更准确捕捉数据生成中的复杂依赖关系。在理论导出的最小性条件下,学习到的表示可等价于真实数据生成过程的潜在变量。实证上,将这些约束应用于主流文生图扩散模型,成功提取其内在的层次化概念图,揭示了模型内部知识组织的新视角。此外,这些基于因果的语义概念可作为细粒度模型引导的控制杠杆,推动透明、可靠的系统发展。
原文摘要 · Abstract (English)
Deep generative models, while revolutionizing fields like image and text generation, largely operate as opaque ``black boxes'', hindering human understanding, control, and alignment. While methods like sparse autoencoders (SAEs) show remarkable empirical success, they often lack theoretical guarantees, risking subjective insights. Our primary objective is to establish a principled foundation for interpretable generative models. We demonstrate that the principle of causal minimality -- favoring the simplest causal explanation -- can endow the latent representations of modern generative models with clear causal interpretation and robust, component-wise identifiable control. We introduce a novel theoretical framework for hierarchical selection models, where higher-level concepts emerge from the constrained composition of lower-level variables, better capturing the complex dependencies in data generation. Under theoretically derived minimality conditions, we show that learned representations can be equivalent to the true latent variables of the data-generating process. Empirically, applying these constraints to leading text-to-image diffusion models allows us to extract their innate hierarchical concept graphs, offering fresh insights into their internal knowledge organization. Furthermore, these causally grounded concepts serve as levers for fine-grained model steering, paving the way for transparent, reliable systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。