arXiv:2409.17527cs.CL2024-09

通过分析模型输出自动推断预训练数据比例,提升大模型数据管理效率。

Data Proportion Detection for Optimized Data Management for Large Language Models

  • 基于模型生成结果反推预训练数据比例,无需原始数据
  • 提出理论证明与算法,初步实验验证方法有效性
  • 适合关注大模型数据优化与可解释性的研究者

大语言模型(LLMs)在众多任务和领域中表现出色,而数据准备在其成功中起关键作用。预训练数据通常融合多个领域的信息,为实现跨领域性能最优,确定最佳数据比例至关重要。然而,当前最先进的大模型很少披露其预训练数据细节,使研究人员难以确定理想的数据比例。本文提出新课题——数据比例检测,可通过分析大模型生成的输出,自动估计其预训练数据的比例。我们提供了严格的理论证明、实用算法及初步实验结果。基于这些发现,我们对有效数据比例检测与数据管理的挑战与未来方向提供了重要见解。

原文摘要 · Abstract (English)

Large language models (LLMs) have demonstrated exceptional performance across a wide range of tasks and domains, with data preparation playing a critical role in achieving these results. Pre-training data typically combines information from multiple domains. To maximize performance when integrating data from various domains, determining the optimal data proportion is essential. However, state-of-the-art (SOTA) LLMs rarely disclose details about their pre-training data, making it difficult for researchers to identify ideal data proportions. In this paper, we introduce a new topic, \textit{data proportion detection}, which enables the automatic estimation of pre-training data proportions by analyzing the generated outputs of LLMs. We provide rigorous theoretical proofs, practical algorithms, and preliminary experimental results for data proportion detection. Based on these findings, we offer valuable insights into the challenges and future directions for effective data proportion detection and data management.

大模型数据管理比例检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。