解决多源异构数据融合难题,提升预测模型性能
Representation Retrieval Learning for Heterogeneous Data Integration
- 构建可检索的表征字典,结合特定数据源的稀疏学习模型
- 提出选择性整合惩罚,显著降低预测误差风险
- 适用于存在变量偏移、缺失数据的复杂真实场景
在大数据时代,大规模、多源、多模态数据日益普遍,为预测建模和科学发现提供了前所未有的机遇。然而,这些数据常表现出复杂的异质性,如协变量偏移、后验漂移和分块缺失,导致现有监督学习算法性能下降。为此,我们提出一种新的表示检索(R2)框架,将表征学习模块字典(代表字典)与数据源特定的稀疏诱导机器学习模型(学习器)相结合。在R2框架下,我们引入每个代表的可整合性概念,并提出一种新颖的选择性整合惩罚(SIP),以显式鼓励更具整合性的代表,从而提升预测性能。理论上,我们证明了R2框架的过量风险界由代表的可整合性刻画,且SIP能有效降低该风险。大量模拟研究验证了R2框架的优越性能及SIP的作用。我们进一步将方法应用于两个真实数据集,证实其经验有效性。
原文摘要 · Abstract (English)
In the era of big data, large-scale, multi-source, multi-modality datasets are increasingly ubiquitous, offering unprecedented opportunities for predictive modeling and scientific discovery. However, these datasets often exhibit complex heterogeneity, such as covariates shift, posterior drift, and blockwise missingness, which worsen predictive performance of existing supervised learning algorithms. To address these challenges simultaneously, we propose a novel Representation Retrieval (R2) framework, which integrates a dictionary of representation learning modules (representer dictionary) with data source-specific sparsity-induced machine learning model (learners). Under the R2 framework, we introduce the notion of integrativeness for each representer, and propose a novel Selective Integration Penalty (SIP) to explicitly encourage more integrative representers to improve predictive performance. Theoretically, we show that the excess risk bound of the R2 framework is characterized by the integrativeness of representers, and SIP effectively improves the excess risk. Extensive simulation studies validate the superior performance of R2 framework and the effect of SIP. We further apply our method to two real-world datasets to confirm its empirical success.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。