arXiv:2409.00284cs.LGcs.AI2024-09

用可生成性评估数据价值,让大模型自己判断数据是否值得保留。

Reframing Data Value for Large Language Models Through the Lens of Plausibility

  • 基于数据能否被模型合理生成来衡量其价值
  • 新方法无需训练即可评估,计算更高效
  • 适合研究数据筛选与模型效率优化的人看

数据估值旨在回答‘这些数据值多少钱’这一核心问题。现有方法多聚焦于判别模型,主要从训练中的实用性来评估数据价值。但随着语言模型规模不断增大,依赖训练的估值方法成本高昂且依赖具体技术。本文提出一种面向大语言模型的数据价值新视角:以数据的可生成性(plausibility)为核心。我们认为,若数据能被模型自身合理生成,则其价值较低。基于符合直觉的评判标准,我们构建了一个从第一性原理出发、计算可操作且具备可证明性质的新价值函数,并在多个场景和数据集上进行了理论分析与评估。

原文摘要 · Abstract (English)

Data valuation seeks to answer the important question, "How much is this data worth?" Existing data valuation methods have largely focused on discriminative models, primarily examining data value through the lens of its utility in training. However, with the push for ever-larger language models, relying on valuation methods that require training becomes increasingly expensive and dependent on specific techniques. We propose an alternative perspective on the data value problem for language models, centering around the plausibility of the data. We posit that data holds lesser value if it can be plausibly generated by the model itself. Starting from some intuitive criteria that align with our notions of valuable data, we develop a novel value function that is computationally tractable and derived from first principles with provable properties. We conduct a theoretical analysis of our value function and evaluate it across multiple scenarios and datasets.

数据估值大模型可生成性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。