arXiv:2505.00903cs.LGcs.CL2025-05NAACL

NeMo-Inspector可高效分析合成数据,提升LLM训练质量

NeMo-Inspector: A Visualization Tool for LLM Generation Analysis

  • 集成推理能力的可视化工具,支持大规模合成数据审查
  • 使低质样本率从46.99%降至19.51%,提升模型准确率1.92%~4.17%
  • 适合需要优化合成数据质量的LLM开发者与研究人员

将大语言模型(LLMs)适配到新任务或提升其整体能力通常需要大量高质量训练数据。当真实数据稀缺或难以获取时,大规模生成的合成数据成为有价值的替代方案。然而,确保合成数据集的质量具有挑战性,开发者需手动检查并修正大量样本以发现错误和改进点,这一过程耗时且依赖专业工具。我们提出NeMo-Inspector,一个开源工具,旨在简化合成数据集的分析,并集成推理功能。通过两个实际案例验证其有效性:使用NeMo-Inspector分析并清洗GSM-Plus合成数据集后,低质量样本比例从46.99%显著下降至19.51%;该工具还帮助识别并修正了OpenMath模型的生成错误,在MATH数据集上使准确率提升1.92%,在GSM8K数据集上提升4.17%,针对基于Nemotron-4-340B生成的合成数据微调的Meta-Llama-3-8B模型。

原文摘要 · Abstract (English)

Adapting Large Language Models (LLMs) to novel tasks and enhancing their overall capabilities often requires large, high-quality training datasets. Synthetic data, generated at scale, serves a valuable alternative when real-world data is scarce or difficult to obtain. However, ensuring the quality of synthetic datasets is challenging, as developers must manually inspect and refine numerous samples to identify errors and areas for improvement. This process is time-consuming and requires specialized tools. We introduce NeMo-Inspector, an open-source tool designed to simplify the analysis of synthetic datasets with integrated inference capabilities. We demonstrate its effectiveness through two real-world cases. Analysis and cleaning of the synthetically generated GSM-Plus dataset with NeMo-Inspector led to a significant decrease in low-quality samples from 46.99% to 19.51%. The tool also helped identify and correct generation errors in OpenMath models, improving accuracy by 1.92% on the MATH dataset and by 4.17% on the GSM8K dataset for a Meta-Llama-3-8B model fine-tuned on synthetic data generated from Nemotron-4-340B.

LLM数据清洗可视化合成数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。