数据质量比数量对小模型性能影响更大,低重复率提升准确率。
Is Training Data Quality or Quantity More Impactful to Small Language Model Performance?
- 对比不同数据量和重复率,评估小模型性能变化
- 25%重复率时准确率升0.87%,100%重复率时准确率降40%
- 研究结果助力降低训练成本,推动AI普惠与可持续发展
本研究通过实证分析探讨了训练数据质量与数量对小型语言模型(SLMs)性能的相对影响,使用TinyStories数据集进行实验。分析了数据集在规模(原大小的25%和50%)与重复率(25%、50%、75%、100%)方面的变化。模型性能基于验证损失、准确率和困惑度进行评估。结果显示,数据质量对SLM整体表现具有更显著的影响,尤其在本实验规模下。适度重复(25%)可提升准确率0.87%,且困惑度仅上升0.52%;但过度重复(100%)导致准确率下降40%。该研究意义不仅限于模型性能,还在于大规模训练带来的巨大经济与计算成本,以及高能耗引发的环境问题。理解数据质量与数量的相对重要性,有助于降低门槛,使先进模型更易获取、更具可持续性。
原文摘要 · Abstract (English)
This study investigates the relative impact of training data quality versus quantity on the performance of small language models (SLMs), utilizing the TinyStories dataset for empirical analysis. Analysis of dataset variations with respect to size (25% and 50% of the original size) and duplication (controlled rates of 25%, 50%, 75%, and 100%) were performed. Model performance was evaluated based on the validation loss, accuracy, and perplexity metrics. Results indicate training data quality plays a more significant role in the overall performance of SLMs, especially given scale of this experiment. Minimal duplication positively impacted model accuracy (+0.87% increase in accuracy at 25% duplication) without significantly increasing perplexity (+0.52% increase going from 0% to 25% duplication) but excessive duplication led to pronounced performance degradation (-40% drop in accuracy at 100% duplication). The implications of this exploration extend beyond just model performance; training large-scale models imposes significant financial and computational burdens, which can be prohibitive for organizations, individuals, and the public at large, especially in developing countries. Additionally, the energy consumption associated with large-scale training raises environmental concerns. Understanding the relative importance of data quality versus quantity could democratize AI technology, making advanced models more accessible and sustainable for all.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。