arXiv:2503.07329cs.CLcs.AI2025-03中稿 · IJCNLP 2025被引 14

随机种子对大模型微调性能影响显著,需重视其稳定性。

Assessing the Macro and Micro Effects of Random Seeds on Fine-Tuning Large Language Models

  • 用GLUE和SuperGLUE基准测试种子影响,量化性能波动
  • 发现宏观准确率与微观预测一致性均有明显差异
  • 适合关注模型可复现性与评估可靠性的研究者

随机种子在大语言模型微调中的影响长期被忽视,尽管其可能显著影响模型表现。本研究基于GLUE和SuperGLUE基准,系统评估了随机种子的作用。通过传统指标(如准确率、F1)计算均值与方差,分析宏观层面的影响;引入新指标‘一致性’,衡量单个预测在多次运行中的稳定性,捕捉微观层面的波动。实验结果表明,随机种子在宏观与微观层面均导致显著性能差异,凸显了微调与评估中需审慎对待随机种子的重要性。

原文摘要 · Abstract (English)

The impact of random seeds in fine-tuning large language models (LLMs) has been largely overlooked despite its potential influence on model performance.In this study, we systematically evaluate the effects of random seeds on LLMs using the GLUE and SuperGLUE benchmarks. We analyze the macro-level impact through traditional metrics like accuracy and F1, calculating their mean and variance to quantify performance fluctuations. To capture the micro-level effects, we introduce a novel metric, consistency, measuring the stability of individual predictions across runs. Our experiments reveal significant variance at both macro and micro levels, underscoring the need for careful consideration of random seeds in fine-tuning and evaluation.

大模型微调随机种子可复现性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。