arXiv:2505.24768cs.CL2025-05AAAI被引 3

分析了大模型微调中数据集多样性对性能的影响,发现响应层微观多样性最重要。

From Macro to Micro: Probing Dataset Diversity in Language Model Fine-Tuning

  • 从宏观、中观到微观分层分析指令与响应的多样性控制策略
  • 在10,000样本固定规模下,响应层微观多样性提升性能最显著
  • 为构建高性能微调数据集提供可操作的多样性设计依据

数据集多样性对机器学习模型的成功训练至关重要,尤其在大语言模型(LLM)的监督微调(SFT)阶段。尽管其重要性日益被认可,但系统性分析仍不足。本文提出一个系统的多样性控制策略分类体系,聚焦于指令组件,在宏观(完整指令语义)和中观(指令单元)层面展开,并首次引入对响应组件的微观多样性分析,具体考察SFT样本中词元的统计分布。实验中,从11.7万条开源SFT样本中构建每组10,000样本的数据集,应用六种涵盖宏观、中观和微观层级的多样性控制策略,分别作用于指令与响应。在这些数据集上微调LLM后发现:宏观与中观策略随多样性增加性能提升,而响应层的微观策略不仅与模型性能相关性更强,且在最大多样性下表现最优。结果为构建高性能SFT数据集提供了切实可行的指导。

原文摘要 · Abstract (English)

Dataset diversity plays a pivotal role for the successful training of many machine learning models, particularly in the supervised fine-tuning (SFT) stage of large language model (LLM) development. Despite increasing recognition of its importance, systematic analyses of dataset diversity still remain underexplored. To address this gap, this work presents a systematic taxonomy of existing diversity-control strategies, which primarily focus on the instruction component, operating at either macroscopic (entire instruction semantics) or mesoscopic levels (instruction units), and furthermore introduces a novel analysis of microscopic diversity within the response component, specifically analyzing the statistical distribution of tokens in SFT training samples. In the experimental evaluation, we construct fixed-size datasets (e.g., 10,000 samples each) from a corpus of 117,000 open-source SFT samples, incorporating six distinct diversity-control strategies spanning macro-, meso-, and microscopic levels applied to both instructions and responses. We then fine-tune LLMs on these datasets to assess the six diversity-control strategies. Results reveal that while macroscopic and mesoscopic strategies lead to higher performance with increasing diversity, the microscopic strategy in responses exhibits both a stronger correlation between model performance and the degree of diversity and superior performance with maximum diversity across all strategies. These findings offer actionable insights for constructing high-performance SFT datasets.

大模型微调数据多样性SFT

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。