arXiv:2510.22567stat.MLcs.LG2025-10被引 3

从因果视角构建通用框架,用无标签数据生成合成标签数据提升模型精度

Semi-Supervised Learning under General Causal Models

  • 基于灵活因果结构设计可学习的因果生成模型
  • 利用无标签数据生成合成标签数据,提升预测性能
  • 适用于标签驱动特征的复杂现实场景,适合因果建模研究者

半监督学习(SSL)旨在同时使用标注和未标注数据训练机器学习模型。尽管未标注数据已被用于多种方式以提高预测精度,但其为何能带来帮助尚不明确。一个有前景的方向是从因果视角理解SSL:根据独立因果机制原则,当标签导致特征而非相反时,未标注数据才可能有效。然而,在实际应用中,特征与标签之间的因果关系往往复杂。本文提出一种适用于一般因果模型的半监督学习框架,其中变量具有灵活的因果关系。我们探索了因果图结构,并设计相应的因果生成模型,这些模型可借助未标注数据进行学习。学习得到的因果生成模型可用于生成合成标注数据,进而训练更准确的预测模型。通过在模拟数据和真实数据上的实证研究,验证了所提方法的有效性。

原文摘要 · Abstract (English)

Semi-supervised learning (SSL) aims to train a machine learning model using both labelled and unlabelled data. While the unlabelled data have been used in various ways to improve the prediction accuracy, the reason why unlabelled data could help is not fully understood. One interesting and promising direction is to understand SSL from a causal perspective. In light of the independent causal mechanisms principle, the unlabelled data can be helpful when the label causes the features but not vice versa. However, the causal relations between the features and labels can be complex in real world applications. In this paper, we propose a SSL framework that works with general causal models in which the variables have flexible causal relations. More specifically, we explore the causal graph structures and design corresponding causal generative models which can be learned with the help of unlabelled data. The learned causal generative model can generate synthetic labelled data for training a more accurate predictive model. We verify the effectiveness of our proposed method by empirical studies on both simulated and real data.

半监督学习因果建模生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。