arXiv:2508.02039cs.LGstat.ML2025-08

无需源数据,通过复用预训练模型实现高效迁移学习

Model Recycling Framework for Multi-Source Data-Free Supervised Transfer Learning

  • 提出模型复用框架,自动筛选相关源模型用于迁移
  • 支持白盒与黑盒场景,实现参数高效训练
  • 适合MaaS服务商构建可复用的预训练模型库

数据隐私问题及源数据获取困难推动了无源数据迁移学习的发展,即仅依赖预训练模型而非原始数据进行训练。该设定带来诸多挑战,因多数现有方法依赖源数据,难以直接应用于无源数据场景。此外,实际应用中还面临如何在无源数据信息下高效选择模型、以及无法完全访问源模型等问题。为此,本文提出一种模型复用框架,可在白盒与黑盒设置下,识别相关源模型子集并加以重用,实现参数高效的模型训练。该框架使模型即服务(MaaS)提供商能够构建高效预训练模型库,为多源无数据监督迁移学习提供可能。

原文摘要 · Abstract (English)

Increasing concerns for data privacy and other difficulties associated with retrieving source data for model training have created the need for source-free transfer learning, in which one only has access to pre-trained models instead of data from the original source domains. This setting introduces many challenges, as many existing transfer learning methods typically rely on access to source data, which limits their direct applicability to scenarios where source data is unavailable. Further, practical concerns make it more difficult, for instance efficiently selecting models for transfer without information on source data, and transferring without full access to the source models. So motivated, we propose a model recycling framework for parameter-efficient training of models that identifies subsets of related source models to reuse in both white-box and black-box settings. Consequently, our framework makes it possible for Model as a Service (MaaS) providers to build libraries of efficient pre-trained models, thus creating an opportunity for multi-source data-free supervised transfer learning.

迁移学习模型复用数据隐私MaaS

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。