用数学方法统一解释大模型微调参数的修改效果,让优化更可控。
A Unified View of Delta Parameter Editing in Post-Trained Large-Scale Models
- 基于损失函数的黎曼求和视角,统一看待参数编辑操作
- 实验验证在多个视觉与语言模型上有效,覆盖ViT、LLaMA 3等
- 改进现有方法如DARE和BitDelta,提升参数编辑通用性
后训练已成为将大规模预训练模型适配到各类任务的关键范式,其效果完全体现在增量参数(即后训练与预训练参数的差异)上。尽管已有大量研究通过剪枝、量化、低秩近似和外推等操作探索增量参数特性,但缺乏系统性分析框架。本文提出一种基于损失函数黎曼求和近似的全新视角,阐明增量参数编辑操作的本质。我们的分析将现有方法按编辑后性能分为三类:表现持平、下降和提升,并解释其如何由黎曼求和项体现,以及如何影响模型性能。在视觉与语言模型(包括ViT、LLaMA 3、Qwen 2、Mistral)上的广泛实验验证了理论发现。此外,我们对DARE和BitDelta等现有技术进行拓展,揭示其在利用增量参数特性方面的局限性,并将其重新组织为通用表达式,以增强后训练模型中增量参数编辑的适用性与有效性。
原文摘要 · Abstract (English)
Post-training has emerged as a crucial paradigm for adapting large-scale pre-trained models to various tasks, whose effects are fully reflected by delta parameters (i.e., the disparity between post-trained and pre-trained parameters). While numerous studies have explored delta parameter properties via operations like pruning, quantization, low-rank approximation, and extrapolation, a unified framework for systematically examining these characteristics has been lacking. In this paper, we propose a novel perspective based on Riemann sum approximation of the loss function to elucidate delta parameter editing operations. Our analysis categorizes existing methods into three classes based on their post-editing performance: competitive, decreased, and improved, explaining how they are expressed by the Riemann sum approximation term and how they alter the model performance. Extensive experiments on both visual and language models, including ViT, LLaMA 3, Qwen 2, and Mistral, corroborate our theoretical findings. Furthermore, we introduce extensions to existing techniques like DARE and BitDelta, highlighting their limitations in leveraging the properties of delta parameters and reorganizing them into general expressions to enhance the applicability and effectiveness of delta parameter editing in post-trained models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。