arXiv:2601.02322stat.MEcs.LG2026-01

根据环境变化动态选择因果或伪相关特征,提升分布外预测性能

Environment-Adaptive Covariate Selection: Learning When to Use Spurious Correlations for Out-of-Distribution Prediction

  • 通过环境级统计量动态决定使用哪些特征
  • 在多种分布外场景下优于固定规则的基线方法
  • 适合需要应对复杂环境变化的下游预测任务

分布外预测中常限制模型仅使用因果或不变特征以避免伪相关。然而当仅观测到部分因果变量时,非因果特征可作为未观测因果变量的代理,在代理关系稳定时提升预测效果,但在环境变化破坏该关系时反而有害。因此最优特征集依赖于具体变化类型。由于不同变化会在无标签特征分布中留下痕迹,本文提出一种环境自适应特征选择算法,将环境级摘要映射为特定环境的特征集合。这些摘要可人工设计或从多环境数据中学习,且可融入先验因果知识作为约束。在模拟和真实数据集上的实验表明,该方法在多种分布外变化下均优于静态因果、不变性及其他非自适应规则。

原文摘要 · Abstract (English)

A common approach to out-of-distribution prediction restricts models to causal or invariant covariates to avoid spurious associations that may change across environments. Despite its theoretical appeal, this strategy can underperform empirical risk minimization when only a subset of the causal parents of the outcome is observed. In such settings, non-causal covariates can serve as proxies for unobserved causal parents and improve prediction when the proxy relationship is stable, but they can hurt when shifts disrupt that relationship. Thus, the optimal covariate set can depend on the specific shift encountered. Because different shifts leave signatures in the unlabeled covariate distribution, we propose an environment-adaptive covariate selection algorithm that maps environment-level summaries to environment-specific covariate sets. These summaries may be hand-crafted or learned from multi-environment data, and prior causal knowledge can be incorporated as constraints. Across simulations and applied datasets, the proposed method improves over static causal, invariant, and other non-adaptive rules under diverse shifts.

分布外预测特征选择因果学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。