arXiv:2412.14809cs.CL2024-12NAACL被引 4

通过参数共振分析,精准筛选大模型训练用的合成数据。

ResoFilter: Fine-grained Synthetic Data Filtering for Large Language Models through Data-Parameter Resonance Analysis

  • 利用微调过程提取数据-参数特征,实现细粒度数据筛选。
  • 仅用一半数据即达全量数据微调效果,数学任务表现优异。
  • 方法通用性强,适合提升合成数据质量与模型训练效率。

大语言模型在多个领域表现出色,基于GPT的合成数据生成已成为主流数据增强方法。然而,增广数据的质量和实用性仍存疑,现有方法缺乏明确的数据特性评估指标。为此,我们提出ResoFilter,一种将模型、数据与任务融合的数据集优化方法。ResoFilter通过微调过程获取数据-参数特征用于数据选择,以模型权重表征数据特性,提升可解释性。实验表明,该方法在数学任务中仅需一半数据即可达到全量微调效果,并在不同模型与领域间展现强泛化能力。该研究为构建高质量合成数据集和评估数据质量提供了新思路,有望显著提升数据增强技术与大模型训练数据质量。代码与数据将在论文接受后公开。

原文摘要 · Abstract (English)

Large language models (LLMs) have shown remarkable effectiveness across various domains, with data augmentation methods utilizing GPT for synthetic data generation becoming prevalent. However, the quality and utility of augmented data remain questionable, and current methods lack clear metrics for evaluating data characteristics. To address these challenges, we propose ResoFilter, a novel method that integrates models, data, and tasks to refine datasets. ResoFilter leverages the fine-tuning process to obtain Data-Parameter features for data selection, offering improved interpretability by representing data characteristics through model weights. Our experiments demonstrate that ResoFilter achieves comparable results to full-scale fine-tuning using only half the data in mathematical tasks and exhibits strong generalization across different models and domains. This method provides valuable insights for constructing synthetic datasets and evaluating high-quality data, offering a promising solution for enhancing data augmentation techniques and improving training dataset quality for LLMs. For reproducibility, we will release our code and data upon acceptance.

大模型数据筛选合成数据微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。