arXiv:2503.14572cs.LGcs.AI2025-03

提出新映射框架,让模型迁移学习性能提升4%。

Robust Weight Imprinting: Insights from Neural Collapse and Proxy-Based Aggregation

  • 用多代理生成与归一化优化映射过程
  • 新方法在迁移任务上比之前高4%
  • 首次关联神经坍缩现象指导代理聚类

基础模型可应用于未见任务,其适应过程称为迁移学习。一种无需参数优化的高效迁移学习方法是映射(imprinting)。本文系统研究了现有映射方法的概念差异,提出通用的\texttt{IMPRINT}框架,包含生成、归一化和聚合三个核心组件。通过该框架深入分析现有方法,发现多代理生成与恰当归一化至关重要。基于此,我们提出一种新变体,利用神经坍缩现象启发的聚类确定代理,使迁移学习性能相较以往提升4%。代码已公开于https://github.com/DATEXIS/IMPRINT。

原文摘要 · Abstract (English)

The capacity of foundation models allows for their application to new, unseen tasks. The adaptation to such tasks is called transfer learning. An efficient transfer learning method that circumvents parameter optimization is imprinting. The conceptual differences between studies on imprinting form the basis of our systematic investigation. In this work, we propose the general \texttt{IMPRINT} framework, identifying three main components: generation, normalization, and aggregation. Through the lens of this framework, we conduct an in-depth analysis and a comparison of the existing methods. Our findings reveal the benefits of representing novel data with multiple proxies in the generation step and show the importance of proper normalization. Beyond an extensive analytical grounding, our framework enables us to propose a novel variant of imprinting which outperforms previous work on transfer learning tasks by 4\%. This variant determines proxies through clustering motivated by the neural collapse phenomenon -- a connection that we draw for the first time. We publicly release our code at https://github.com/DATEXIS/IMPRINT.

迁移学习映射方法神经坍缩

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。