arXiv:2410.12178cs.LGstat.ML2024-10EMNLP被引 18

通过平衡各层训练质量,提升小数据下的模型表现

Model Balancing Helps Low-data Training and Fine-tuning

  • 按层调整学习率,解决不同层训练不均衡问题
  • 数据越少,性能提升越明显,小样本下效果更优
  • 可作为通用插件方法,适用于NLP与科学机器学习

近期基础模型的发展强调了使用少量精心筛选的数据集对预训练模型进行领域对齐的重要性。相关研究凸显了低数据训练与微调的关键作用。这一主题在自然语言处理(NLP)中已广为人知,如今在科学机器学习(SciML)领域也日益受到关注。为应对低数据训练与微调的局限性,本文借鉴重尾自正则化(HT-SR)理论,分析经验谱密度(ESDs)形状,揭示了模型各层间训练质量的不平衡现象。为此,我们采用一种新提出的逐层学习率调度器TempBalance,有效平衡各层训练质量,显著提升NLP与SciML任务中的低数据训练与微调性能。值得注意的是,随着可用微调数据减少,TempBalance带来的性能增益持续增大。对比分析进一步证明了其有效性及作为‘即插即用’方法的强适应性。

原文摘要 · Abstract (English)

Recent advances in foundation models have emphasized the need to align pre-trained models with specialized domains using small, curated datasets. Studies on these foundation models underscore the importance of low-data training and fine-tuning. This topic, well-known in natural language processing (NLP), has also gained increasing attention in the emerging field of scientific machine learning (SciML). To address the limitations of low-data training and fine-tuning, we draw inspiration from Heavy-Tailed Self-Regularization (HT-SR) theory, analyzing the shape of empirical spectral densities (ESDs) and revealing an imbalance in training quality across different model layers. To mitigate this issue, we adapt a recently proposed layer-wise learning rate scheduler, TempBalance, which effectively balances training quality across layers and enhances low-data training and fine-tuning for both NLP and SciML tasks. Notably, TempBalance demonstrates increasing performance gains as the amount of available tuning data decreases. Comparative analyses further highlight the effectiveness of TempBalance and its adaptability as an "add-on" method for improving model performance.

低数据训练模型平衡微调优化层间均衡

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。