arXiv:2503.04807cs.CLcs.AI2025-03ACL被引 1

提醒研究者:指令微调数据质量评估需规范超参数设置,否则结论不可靠。

Call for Rigor in Reporting Quality of Instruction Tuning Data

  • 用模型表现反推数据质量时,超参数选择应有依据,不能随意设定。
  • 实验显示,不同超参数下同一数据集可得出完全相反的结论。
  • 适合关注大模型训练严谨性的研究者和审稿人阅读。

指令微调对大语言模型与用户意图对齐至关重要。现有研究普遍认为指令微调数据质量与模型对齐性能强相关,通常通过模型训练表现来评估数据质量。然而我们发现,多数研究中模型训练的超参数选择缺乏充分依据,不同研究间即使使用相同模型和数据,也存在显著差异。本文通过在LIMA数据集及1,000个Alpaca数据点上的实验表明,任意设定超参数可能导致任意结论,揭示了当前评估方法的潜在风险,强调必须严谨对待数据质量验证过程。

原文摘要 · Abstract (English)

Instruction tuning is crucial for adapting large language models (LLMs) to align with user intentions. Numerous studies emphasize the significance of the quality of instruction tuning (IT) data, revealing a strong correlation between IT data quality and the alignment performance of LLMs. In these studies, the quality of IT data is typically assessed by evaluating the performance of LLMs trained with that data. However, we identified a prevalent issue in such practice: hyperparameters for training models are often selected arbitrarily without adequate justification. We observed significant variations in hyperparameters applied across different studies, even when training the same model with the same data. In this study, we demonstrate the potential problems arising from this practice and emphasize the need for careful consideration in verifying data quality. Through our experiments on the quality of LIMA data and a selected set of 1,000 Alpaca data points, we demonstrate that arbitrary hyperparameter decisions can make any arbitrary conclusion.

指令微调模型评估实验严谨性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。