评测大模型生成合成数据的能力,发现不同模型各有优劣。
Evaluating Language Models as Synthetic Data Generators
- 构建统一基准AgoraBench,系统评估6个大模型生成能力
- GPT-4o擅长创造新问题,Claude-3.5-Sonnet更善优化旧问题
- 生成质量与解题能力无关,需综合响应质量、困惑度等指标
随着合成数据在语言模型后训练中的广泛应用,模型生成高质量数据的能力已与直接解决问题的能力同样重要。现有研究虽聚焦高效生成方法,却缺乏在统一设置下对不同语言模型作为数据生成器的系统比较。为此,我们提出AgoraBench基准,提供标准化设置与评估指标。通过使用6个语言模型生成126万条训练样本,并训练99个学生模型,我们揭示了若干关键洞见:首先,不同模型表现出显著差异——例如,GPT-4o在生成新问题方面表现突出,而Claude-3.5-Sonnet在增强已有问题上更优;其次,数据生成能力与模型解题能力无必然关联,响应质量、困惑度及指令难度等内在特征共同构成更优预测指标;最后,输出格式选择与成本敏感的模型选型对生成效果有显著影响。
原文摘要 · Abstract (English)
Given the increasing use of synthetic data in language model (LM) post-training, an LM's ability to generate high-quality data has become nearly as crucial as its ability to solve problems directly. While prior works have focused on developing effective data generation methods, they lack systematic comparison of different LMs as data generators in a unified setting. To address this gap, we propose AgoraBench, a benchmark that provides standardized settings and metrics to evaluate LMs' data generation abilities. Through synthesizing 1.26 million training instances using 6 LMs and training 99 student models, we uncover key insights about LMs' data generation capabilities. First, we observe that LMs exhibit distinct strengths. For instance, GPT-4o excels at generating new problems, while Claude-3.5-Sonnet performs better at enhancing existing ones. Furthermore, our analysis reveals that an LM's data generation ability doesn't necessarily correlate with its problem-solving ability. Instead, multiple intrinsic features of data quality-including response quality, perplexity, and instruction difficulty-collectively serve as better indicators. Finally, we demonstrate that strategic choices in output format and cost-conscious model selection significantly impact data generation effectiveness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。