arXiv:2411.19923cs.LGstat.ML2024-11被引 5

在未知混杂因素下实现高效分布外鲁棒性,仅用一个额外变量即可。

Scalable Out-of-distribution Robustness in the Presence of Unobserved Confounders

  • 基于单一额外变量的可识别性假设,构建简单鲁棒预测器。
  • 在多个基准任务上表现优于现有复杂方法。
  • 适用于测试分布与训练不一致且无测试数据的情况,适合工业级部署。

我们研究因未观测混杂因子 $Z$ 导致的分布外(OOD)泛化问题,该因子同时影响特征 $X$ 与标签 $Y$,导致预测分布具有异质性:$P(Y | X) = E_{P(Z | X)}[P(Y | X,Z)]$,使传统协变量或标签分布偏移假设失效。与传统领域自适应不同,本任务不假设训练时可获取测试特征分布 $X^ ext{te}$。这带来四大挑战:(a) 训练时 $Z^ ext{tr}$ 未被观测,(b) $P^ ext{te}(Z) \neq P^ ext{tr}(Z)$,(c) 训练时不可见 $X^ ext{te}$,(d) 预测分布依赖于 $P^ ext{te}(Z)$。已有工作需多个附加变量才能识别潜在分布,而本文提出一组更宽松的可识别性假设,仅需一个额外变量即可实现有效估计。实验表明该方法在多个基准任务上取得优异性能。

原文摘要 · Abstract (English)

We consider the task of out-of-distribution (OOD) generalization, where the distribution shift is due to an unobserved confounder ($Z$) affecting both the covariates ($X$) and the labels ($Y$). This confounding introduces heterogeneity in the predictor, i.e., $P(Y | X) = E_{P(Z | X)}[P(Y | X,Z)]$, making traditional covariate and label shift assumptions unsuitable. OOD generalization differs from traditional domain adaptation in that it does not assume access to the covariate distribution ($X^\text{te}$) of the test samples during training. These conditions create a challenging scenario for OOD robustness: (a) $Z^\text{tr}$ is an unobserved confounder during training, (b) $P^\text{te}(Z) \neq P^\text{tr}(Z)$, (c) $X^\text{te}$ is unavailable during training, and (d) the predictive distribution depends on $P^\text{te}(Z)$. While prior work has developed complex predictors requiring multiple additional variables for identifiability of the latent distribution, we explore a set of identifiability assumptions that yield a surprisingly simple predictor using only a single additional variable. Our approach demonstrates superior empirical performance on several benchmark tasks.

分布外泛化混杂因子鲁棒学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。