arXiv:2411.00039cs.CLcs.AI2024-11

通过线性链变换提升大模型微调的优化能力

Linear Chain Transformation: Expanding Optimization Dynamics for Fine-Tuning Large Language Models

  • 在微调中引入多层线性变换,扩展参数更新的秩空间
  • 相比现有方法,提升泛化性能并减少可训练参数
  • 适合追求高效微调的大模型应用开发者

大语言模型的微调对于适应特定下游任务至关重要。本文提出线性链变换(LinChain),在微调过程中引入一系列线性变换,以丰富优化动态。通过将多个线性变换融入参数更新过程,LinChain扩展了有效更新秩,增强了模型学习复杂任务特异性表征的能力。实验表明,该方法在多种基准任务上显著优于当前最优方法,提供了更灵活的训练优化路径,同时保持了模型推理效率。结果表明,LinChain能提升泛化性能、减少可训练参数并增强任务适配能力,是一种极具潜力的大模型微调策略。

原文摘要 · Abstract (English)

Fine-tuning large language models (LLMs) has become essential for adapting pretrained models to specific downstream tasks. In this paper, we propose Linear Chain Transformation (LinChain), a novel approach that introduces a sequence of linear transformations during fine-tuning to enrich optimization dynamics. By incorporating multiple linear transformations into the parameter update process, LinChain expands the effective rank of updates and enhances the model's ability to learn complex task-specific representations. We demonstrate that this method significantly improves the performance of LLM fine-tuning over state-of-the-art methods by providing more flexible optimization paths during training, while maintaining the inference efficiency of the resulting model. Our experiments on various benchmark tasks show that LinChain leads to better generalization, fewer learnable parameters, and improved task adaptation, making it a compelling strategy for LLM fine-tuning.

大模型微调优化动态线性变换

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。