arXiv:2502.13514cs.CL2025-02

不用重训模型,几条样本就能评估数据策略效果

Shall Your Data Strategy Work? Perform a Swift Study

  • 用梯度投影分析少量探针样本,快速评估数据策略
  • 验证显示新方法与真实训练结果一致,准确率提升显著
  • 适合想高效优化指令微调数据的研究者和工程师

本文提出一种快速评估指令微调数据策略有效性的方法,仅需少量探针样本即可完成评估,无需重新训练模型。该方法基于梯度数据影响估计,通过分析选定策略下探针样本在评估样本上的梯度投影,判断其优势。基于此,我们开展了三项快速研究,分别考察了思维链(Chain-of-thought, CoT)数据、查询澄清数据和响应评估数据对模型泛化能力的潜力。随后进行验证研究,针对每项策略构建专属训练数据集,并对比使用与不使用该数据集时的模型表现。验证结果与快速研究结论一致,证实了所提方法的有效性。

原文摘要 · Abstract (English)

This work presents a swift method to assess the efficacy of particular types of instruction-tuning data, utilizing just a handful of probe examples and eliminating the need for model retraining. This method employs the idea of gradient-based data influence estimation, analyzing the gradient projections of probe examples from the chosen strategy onto evaluation examples to assess its advantages. Building upon this method, we conducted three swift studies to investigate the potential of Chain-of-thought (CoT) data, query clarification data, and response evaluation data in enhancing model generalization. Subsequently, we embarked on a validation study to corroborate the findings of these swift studies. In this validation study, we developed training datasets tailored to each studied strategy and compared model performance with and without the use of these datasets. The results of the validation study aligned with the findings of the swift studies, validating the efficacy of our proposed method.

数据评估指令微调快速实验梯度分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。