在因果模型中,用信息瓶颈法实现跨域目标变量的精准补全。
Causally-Aware Information Bottleneck for Domain Adaptation
- 基于信息瓶颈构建稳定表征,保留预测目标所需信息。
- 线性场景下解出闭式解,非线性时用变分方法可零样本部署。
- 适用于高维数据,适合需要跨域推断的因果建模任务。
我们研究因果系统中的常见领域自适应设置:目标变量在源域可观测,但在目标域完全缺失。目标是通过剩余可观测变量,在各种分布偏移下对目标变量进行补全。为此,我们将其建模为学习紧凑且机制稳定的表示——既保留预测目标所需的有用信息,又剔除虚假变化。对于线性高斯因果模型,我们推导出闭式高斯信息瓶颈(GIB)解,其形式等价于一种类似典型相关分析(CCA)的投影,并可在需要时融入有向无环图(DAG)先验。对于非线性或非高斯数据,我们引入变分信息瓶颈(VIB)编码器-预测器结构,支持高维扩展,仅需在源域训练即可零样本部署至目标域。在合成与真实数据集上,该方法始终实现准确补全,验证了其在高维因果模型中的实用性,提供了一个统一、轻量的因果领域自适应工具包。
原文摘要 · Abstract (English)
We tackle a common domain adaptation setting in causal systems. In this setting, the target variable is observed in the source domain but is entirely missing in the target domain. We aim to impute the target variable in the target domain from the remaining observed variables under various shifts. We frame this as learning a compact, mechanism-stable representation. This representation preserves information relevant for predicting the target while discarding spurious variation. For linear Gaussian causal models, we derive a closed-form Gaussian Information Bottleneck (GIB) solution. This solution reduces to a canonical correlation analysis (CCA)-style projection and offers Directed Acyclic Graph (DAG)-aware options when desired. For nonlinear or non-Gaussian data, we introduce a Variational Information Bottleneck (VIB) encoder-predictor. This approach scales to high dimensions and can be trained on source data and deployed zero-shot to the target domain. Across synthetic and real datasets, our approach consistently attains accurate imputations, supporting practical use in high-dimensional causal models and furnishing a unified, lightweight toolkit for causal domain adaptation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。