arXiv:2505.22157cs.CL2025-05中稿 · Findings of the As…被引 1

提出高效通用的数据筛选方法,提升大模型微调效果

LASER: Stratified Selective Sampling for Instruction Tuning with Dedicated Scoring Strategy

  • 分层选择+专用评分策略,兼顾效率与泛化能力
  • 通过任务分类控制数据组成,适配多用途模型训练
  • 结合嵌入与聚类增强多样性,适合大规模微调场景

近期研究发现,大型语言模型的后训练数据集可大幅缩减而性能不显著下降。然而,现有数据选择方法往往计算开销高或仅适用于特定领域。本文提出一种多步骤流水线:高效将数据点分组,利用专用模型估计质量,并采用轻量级鲁棒方法评估难度。基于任务的分类可调控最终数据组成,对多用途模型微调至关重要。为保障多样性,改进了先前方法,使用嵌入模型与聚类算法。该集成策略在极低开销下实现高性能微调。

原文摘要 · Abstract (English)

Recent work shows that post-training datasets for LLMs can be substantially downsampled without noticeably deteriorating performance. However, data selection often incurs high computational costs or is limited to narrow domains. In this paper, we demonstrate that data selection can be both -- efficient and universal -- by using a multi-step pipeline in which we efficiently bin data points into groups, estimate quality using specialized models, and score difficulty with a robust, lightweight method. Task-based categorization allows us to control the composition of our final data -- crucial for finetuning multi-purpose models. To guarantee diversity, we improve upon previous work using embedding models and a clustering algorithm. This integrated strategy enables high-performance fine-tuning with minimal overhead.

数据筛选大模型微调高效训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。