用智能评分系统精选数据,小样本反而比大样本更有效。
Improving Data Efficiency via Curating LLM-Driven Rating Systems
- 通过分析评分错误模式,动态修正大模型生成的评分结果。
- 仅3.3%的精选数据超越30万样本的全量数据集。
- 适合追求高效训练、减少冗余数据的研究者使用。
指令微调对大语言模型适配下游任务至关重要,近期研究显示少量人工标注数据可优于大规模数据集,挑战了传统的数据规模定律。尽管基于大模型的数据质量评分系统能降低成本,但即使在GPT-4等强大模型上仍存在误差与偏差。本文提出DS2——一种面向数据选择的多样性感知评分优化方法,通过构建评分转移矩阵系统建模误差模式,校正大模型评分并提升所选样本多样性。实验表明,仅3.3%的经筛选数据集在多个机器对齐基准上表现优于30万样本的完整数据集;且在1000个样本规模下,性能可匹配或超越人类标注数据集LIMA。这些发现挑战了传统数据规模假设,揭示冗余低质样本可能损害性能,再次印证‘多未必好’。
原文摘要 · Abstract (English)
Instruction tuning is critical for adapting large language models (LLMs) to downstream tasks, and recent studies have demonstrated that small amounts of human-curated data can outperform larger datasets, challenging traditional data scaling laws. While LLM-based data quality rating systems offer a cost-effective alternative to human annotation, they often suffer from inaccuracies and biases, even in powerful models like GPT-4. In this work, we introduce DS2, a Diversity-aware Score curation method for Data Selection. By systematically modeling error patterns through a score transition matrix, DS2 corrects LLM-based scores and promotes diversity in the selected data samples. Our approach shows that a curated subset (just 3.3% of the original dataset) outperforms full-scale datasets (300k samples) across various machine-alignment benchmarks, and matches or surpasses human-aligned datasets such as LIMA with the same sample size (1k samples). These findings challenge conventional data scaling assumptions, highlighting that redundant, low-quality samples can degrade performance and reaffirming that "more can be less."
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。