按数据实际贡献定价,让高质量数据值更多钱。
Utility-Aware Data Pricing: Token-Level Quality and Empirical Training Gain for LLMs
- 用熵和质量分评估每个词元的有用性
- 实测训练增益显示新方法排名准确率超基线
- 支持可验证、防篡改的数据交易记录
基于“行数×质量系数”的传统数据估值方法无法捕捉数据对大语言模型能力的非线性贡献。本文提出一种动态数据估值框架,从静态计数转向基于效用的定价。方法包含三层:(1) 使用香农熵和数据质量分计算词元级信息密度;(2) 通过影响函数、代理模型和数据沙普利值实测训练增益;(3) 利用哈希承诺、默克尔树和防篡改训练账本实现密码学可验证性。在指令遵循、数学推理和代码摘要三个真实场景中进行充分实验,结果显示基于代理模型的实测增益与实际效用排名高度一致,显著优于行数和词元数基线。该框架支持公平的数据即服务经济,使高推理价值数据按其真实贡献定价,并保障数据市场的透明与可审计。
原文摘要 · Abstract (English)
Traditional data valuation methods based on ``row-count $\times$ quality coefficient'' paradigms fail to capture the nuanced, nonlinear contributions that data makes to Large Language Model (LLM) capabilities. This paper presents a dynamic data valuation framework that transitions from static accounting to utility-based pricing. Our approach operates on three layers: (1) token-level information density metrics using Shannon entropy and Data Quality Scores; (2) empirical training gain measurement through influence functions, proxy model strategies, and Data Shapley values; and (3) cryptographic verifiability through hash-based commitments, Merkle trees, and a tamper-evident training ledger. We provide comprehensive experimental validation on three real domains (instruction following, mathematical reasoning, and code summarization), demonstrating that proxy-based empirical gain achieves near-perfect ranking alignment with realized utility, substantially outperforming row-count and token-count baselines. This framework enables a fair Data-as-a-Service economy where high-reasoning data is priced according to its actual contribution to model intelligence, while providing the transparency and auditability necessary for trustworthy data markets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。