arXiv:2602.17187stat.MLcs.LG2026-02中稿 · ICML

利用无标签数据提升模型对分布偏移的鲁棒性,无需依赖标注数据。

Anti-causal domain generalization: Leveraging unlabeled data

  • 在反因果框架下,通过无标签数据捕捉环境扰动方向。
  • 针对均值和协方差变化设计正则项,实现最坏情况下的最优性保证。
  • 适用于标注稀缺场景,尤其适合物理系统与生理信号建模。

领域泛化旨在学习在新环境部署时仍能保持性能的预测模型。现有方法通常需要多个训练环境的标注数据,限制了其在标注数据稀缺时的应用。本文研究反因果设置下的领域泛化,其中结果导致可观测特征。在此结构中,影响特征的环境扰动不会传递至结果,因此可通过对模型对这些扰动的敏感性进行正则化来增强鲁棒性。关键的是,估计扰动方向无需标签,从而可利用多环境的无标签数据。本文提出两种方法,分别惩罚模型对特征均值和协方差跨环境变化的敏感性,并证明在特定环境类上具有最坏情况下的最优性。实验在受控物理系统和生理信号数据集上验证了方法的有效性。

原文摘要 · Abstract (English)

The problem of domain generalization concerns learning predictive models that are robust to distribution shifts when deployed in new, previously unseen environments. Existing methods typically require labeled data from multiple training environments, limiting their applicability when labeled data are scarce. In this work, we study domain generalization in an anti-causal setting, where the outcome causes the observed covariates. Under this structure, environment perturbations that affect the covariates do not propagate to the outcome, which motivates regularizing the model's sensitivity to these perturbations. Crucially, estimating these perturbation directions does not require labels, enabling us to leverage unlabeled data from multiple environments. We propose two methods that penalize the model's sensitivity to variations in the mean and covariance of the covariates across environments, respectively, and prove that these methods have worst-case optimality guarantees under certain classes of environments. Finally, we demonstrate the empirical performance of our approach on a controlled physical system and a physiological signal dataset.

领域泛化无监督学习反因果

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。