提出可扩展的第三方大模型数据估值方法,高效评估训练样本价值。
ALinFiK: Learning to Approximate Linearized Future Influence Kernel for Scalable Third-Party LLM Data Valuation
- 用线性化未来影响核衡量数据对模型性能的贡献
- 在参数增长时仍保持高效,优于现有基线方法
- 适合数据提供方与模型开发者共同使用
大型语言模型(LLMs)高度依赖高质量训练数据,因此在有限预算下进行数据估值对优化模型性能至关重要。本文提出一种第三方数据估值方法,兼顾数据提供方与模型开发者的利益。引入线性化未来影响核(LinFiK),用于评估单个数据样本在训练过程中提升模型性能的价值。进一步提出ALinFiK学习策略,以近似计算LinFiK,实现可扩展的数据估值。全面评估表明,该方法在有效性和效率上均超越现有基线,在模型参数增大时展现出显著的可扩展优势。
原文摘要 · Abstract (English)
Large Language Models (LLMs) heavily rely on high-quality training data, making data valuation crucial for optimizing model performance, especially when working within a limited budget. In this work, we aim to offer a third-party data valuation approach that benefits both data providers and model developers. We introduce a linearized future influence kernel (LinFiK), which assesses the value of individual data samples in improving LLM performance during training. We further propose ALinFiK, a learning strategy to approximate LinFiK, enabling scalable data valuation. Our comprehensive evaluations demonstrate that this approach surpasses existing baselines in effectiveness and efficiency, demonstrating significant scalability advantages as LLM parameters increase.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。