arXiv:2511.10919stat.MLcs.LG2025-11

融合多源异构数据,提升无负样本场景下的预测准确率

Heterogeneous Multisource Transfer Learning via Model Averaging for Positive-Unlabeled Data

  • 通过模型平均整合有标签、半监督和无负样本数据
  • 在信用风险数据上显著优于对比方法,尤其在标注数据少时
  • 适用于隐私敏感领域,如金融风控与医疗诊断

正例-无标签(PU)学习因缺乏明确的负样本标注,在欺诈检测和医疗诊断等高风险领域面临独特挑战。为应对数据稀缺与隐私约束,本文提出一种基于模型平均的异构多源迁移学习框架,无需直接共享数据即可融合来自全二值标注、半监督及PU数据集的信息。针对每类源域,构建定制化逻辑回归模型,并通过模型平均将知识迁移到PU目标域。最优组合权重由最小化Kullback-Leibler散度的交叉验证准则确定。建立了权重最优性与收敛性的理论保证,涵盖模型误设与正确设定情形,并进一步扩展至高维设置,采用稀疏惩罚估计器。大量模拟实验与真实信用风险数据分析表明,该方法在预测精度与鲁棒性方面均优于对比方法,尤其在标注数据有限且环境异构条件下表现突出。

原文摘要 · Abstract (English)

Positive-Unlabeled (PU) learning presents unique challenges due to the lack of explicitly labeled negative samples, particularly in high-stakes domains such as fraud detection and medical diagnosis. To address data scarcity and privacy constraints, we propose a novel transfer learning with model averaging framework that integrates information from heterogeneous data sources - including fully binary labeled, semi-supervised, and PU data sets - without direct data sharing. For each source domain type, a tailored logistic regression model is conducted, and knowledge is transferred to the PU target domain through model averaging. Optimal weights for combining source models are determined via a cross-validation criterion that minimizes the Kullback-Leibler divergence. We establish theoretical guarantees for weight optimality and convergence, covering both misspecified and correctly specified target models, with further extensions to high-dimensional settings using sparsity-penalized estimators. Extensive simulations and real-world credit risk data analyses demonstrate that our method outperforms other comparative methods in terms of predictive accuracy and robustness, especially under limited labeled data and heterogeneous environments.

PU学习迁移学习信用风险模型平均

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。