arXiv:2604.26128stat.MLcs.LG2026-04被引 1

显式建模环境差异,提升跨环境预测鲁棒性

Robust Representation Learning through Explicit Environment Modeling

  • 通过显式建模环境变化并消除其影响,学习更鲁棒的表示
  • 在多个挑战性设置中,优于传统因果不变表示方法
  • 适合环境分布差异大、需泛化到未知环境的任务

我们研究从多个环境中收集的带标签数据中学习问题,这些环境的数据分布可能不同。传统方法多从因果视角出发,寻找不变表示以保留因果因子、剔除伪相关。但该方法假设环境对目标变量无直接影响,这一假设在实际中常不成立。本文考虑环境有直接作用的情形,仍希望学习出能在先前未见环境中平均表现良好的表示。为此,我们研究通过显式建模环境间差异,并对差异进行边际化处理所获得的表示。我们分析了这些表示的性质,明确了其相较于因果不变表示方法的优势场景。提出基于广义随机截距模型的方法,该类模型支持有效边际化,且可分析其泛化性能。实验证明,在多种复杂场景下,该方法显著优于现有不变学习方法。

原文摘要 · Abstract (English)

We consider learning from labeled data collected across multiple environments, where the data distribution may vary across these environments. This problem is commonly approached from a causal perspective, seeking invariant representations that retain causal factors while discarding spurious ones. However, this framework assumes that the environment has no direct effect on the target. In contrast, we consider settings in which this assumption fails, but still aim to learn representations that support robust prediction on average across previously unseen environments. To this end, we study representations learned by explicitly modeling variation across environments and then marginalizing that variation out. We analyze the resulting representations and characterize when they are preferable to those learned by causal invariant-representation methods. We propose a concrete method based on generalized random-intercept models, a class of predictors in which such marginalization is possible, and study their generalization properties. Empirically, we show that these models outperform invariant-learning methods across a range of challenging settings.

表示学习鲁棒性环境建模泛化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。