通过辅助标签与不变性原则,分离跨环境共享与特有潜在因子,提升模型泛化能力。
Unveiling Invariant and Transferable Latent Factors Across Heterogeneous Environments via ATLAS
- 基于不变性原理,分离共享因子与环境特异因子。
- 在辅助标签帮助下,实现新环境中近最优的可迁移预测。
- 适合需要跨域泛化的机器学习任务,如医疗、金融建模。
本文研究多环境因子模型,高维协变量来自异质环境,部分环境中提供辅助标签。协变量联合分布随环境变化,而潜在结构分为具有共享载荷的不变因子与具有环境特异载荷的异质因子。该模型源于迁移学习与潜在因子回归,旨在获得稳定低维表示,以支持响应 $Y$ 的解释与鲁棒外样本预测。利用不变性原则,我们证明在最简结构条件下,不变因子与异质因子可被解耦。基于此,提出 ATLAS(辅助标签与不变性引导的跨环境潜在对齐),一种统一方法:利用不变性原则分离对齐的不变因子与未对齐的异质因子,并借助辅助标签从异质因子中提取预测不变且可迁移的因子。ATLAS 在下游潜在因子回归中达到近似最优性能;当有辅助标签时,可通过完整潜在信号实现新环境中的可迁移预测;否则退化为仅依赖不变因子的稳健预测。我们建立了恢复不变与异质因子、识别所有响应不变因子以及估计 $Y$ 中不变信号的精确非渐近误差界。
原文摘要 · Abstract (English)
This paper considers a multi-environment factor model in which high-dimensional covariates are collected from heterogeneous environments, with auxiliary labels available in a subset of these environments. The joint distribution of the covariates may vary across environments, whereas the latent structure is decomposed into invariant factors with shared loadings and heterogeneous factors with environment-specific loadings. Such a model is motivated by transfer learning and latent factor regression, where one seeks stable low-dimensional representations for both interpretation and robust out-of-sample prediction of the response $Y$. Leveraging the invariance principle, we show that the invariant and heterogeneous factors are disentangled under a minimal structural condition. Based on this, we propose ATLAS, an Auxiliary-label and invariance-guided Transfer via Latent Alignment across heterogeneous environmentS. ATLAS is a unified procedure that leverages the invariance principle to separate aligned invariant and unaligned heterogeneous factors, and further exploits supervision from auxiliary labels to extract prediction-invariant and transferable factors from those unaligned heterogeneous factors. ATLAS yields near-oracle performance for downstream latent factor regression, enables transferable prediction in new environments through the full latent signal when auxiliary labels are available, and reduces to robust invariant-factor-only prediction otherwise. We establish sharp non-asymptotic error bounds for recovering invariant and heterogeneous factors, identifying all the response-invariant factors, and estimating the invariant signal in $Y$.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。