arXiv:2601.07247stat.MLcs.LG2026-01

解决缺失数据下的环境不变学习,提升模型鲁棒性与因果解释力

Multi-environment Invariance Learning with Missing Data

  • 基于不变性目标设计缺失数据下的新估计器
  • 理论证明变量选择正确且误差收敛速率受缺失比例影响
  • 适用于数据不全但需稳定预测的现实场景

在领域泛化中,应对分布变化是核心挑战。不变性学习通过识别跨环境稳定的特征,捕捉潜在因果关系,从而提升模型泛化能力。然而,实际中因数据收集成本高,常面临结果数据缺失问题,限制了对环境异质性的充分利用。本文推导了在结果缺失情况下的不变性目标估计器,建立了非渐近的变量选择性质和ℓ₂误差收敛速率理论,其性能受缺失比例及各环境插补模型质量影响。通过大量模拟实验和UCI共享单车数据集上的应用验证,即使使用有偏插补模型,只要偏差在合理范围内,该估计器仍具高效性且预测误差更低。

原文摘要 · Abstract (English)

Learning models that can handle distribution shifts is a key challenge in domain generalization. Invariance learning, an approach that focuses on identifying features invariant across environments, improves model generalization by capturing stable relationships, which may represent causal effects when the data distribution is encoded within a structural equation model (SEM) and satisfies modularity conditions. This has led to a growing body of work that builds on invariance learning, leveraging the inherent heterogeneity across environments to develop methods that provide causal explanations while enhancing robust prediction. However, in many practical scenarios, obtaining complete outcome data from each environment is challenging due to the high cost or complexity of data collection. This limitation in available data hinders the development of models that fully leverage environmental heterogeneity, making it crucial to address missing outcomes to improve both causal insights and robust prediction. In this work, we derive an estimator from the invariance objective under missing outcomes. We establish non-asymptotic guarantees on variable selection property and $\ell_2$ error convergence rates, which are influenced by the proportion of missing data and the quality of imputation models across environments. We evaluate the performance of the new estimator through extensive simulations and demonstrate its application using the UCI Bike Sharing dataset to predict the count of bike rentals. The results show that despite relying on a biased imputation model, the estimator is efficient and achieves lower prediction error, provided the bias is within a reasonable range.

不变性学习缺失数据因果推理领域泛化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。