arXiv:2501.18084stat.MLcs.LG2025-01

无需标签即可融合多个模型,提升新人群预测的准确性与鲁棒性。

U-aggregation: Unsupervised Aggregation of Multiple Learning Algorithms

  • 基于随机矩阵理论,通过方差稳定与稀疏信号迭代恢复融合模型
  • 在无真实标签情况下仍能准确估计个体风险并评估模型表现
  • 适用于遗传风险预测等缺乏标签的真实场景,尤其适合开放模型集成

在开放科学与开源机器学习日益普及的背景下,大量预训练模型可被直接调用,减少标注与训练成本。然而,在目标人群上选择最优模型仍面临挑战,原因包括迁移能力有限、数据异质性以及真实标签难以获取。本文提出一种无监督模型融合方法U-aggregation,可在无观测标签的情况下整合多个预训练模型,以增强新人群中的性能与鲁棒性。该方法不依赖于监督信号,且能处理模型层面和个体层面的异方差性,以及对抗性模型的存在。基于随机矩阵理论,U-aggregation引入方差稳定步骤和迭代稀疏信号恢复机制,提升个体真实风险估计精度,并评估候选模型相对表现。我们通过理论分析与系统数值实验验证其性质,并以PGS Catalog中公开模型为例,展示了其在复杂性状遗传风险预测中的实际应用潜力。

原文摘要 · Abstract (English)

Across various domains, the growing advocacy for open science and open-source machine learning has made an increasing number of models publicly available. These models allow practitioners to integrate them into their own contexts, reducing the need for extensive data labeling, training, and calibration. However, selecting the best model for a specific target population remains challenging due to issues like limited transferability, data heterogeneity, and the difficulty of obtaining true labels or outcomes in real-world settings. In this paper, we propose an unsupervised model aggregation method, U-aggregation, designed to integrate multiple pre-trained models for enhanced and robust performance in new populations. Unlike existing supervised model aggregation or super learner approaches, U-aggregation assumes no observed labels or outcomes in the target population. Our method addresses limitations in existing unsupervised model aggregation techniques by accommodating more realistic settings, including heteroskedasticity at both the model and individual levels, and the presence of adversarial models. Drawing on insights from random matrix theory, U-aggregation incorporates a variance stabilization step and an iterative sparse signal recovery process. These steps improve the estimation of individuals' true underlying risks in the target population and evaluate the relative performance of candidate models. We provide a theoretical investigation and systematic numerical experiments to elucidate the properties of U-aggregation. We demonstrate its potential real-world application by using U-aggregation to enhance genetic risk prediction of complex traits, leveraging publicly available models from the PGS Catalog.

模型融合无监督学习遗传预测开放模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。