提出新方法评估数据集价值,让选数据更科学高效。
How Much Is a Dataset Worth? Scaling Laws, the Vendi Score, and Matrix Spectral Functions

- 用矩阵谱函数统一多种数据评估指标,理论更完整。
- 优化速度提升3.5万倍,可在大规模数据上直接计算评分。
- 发现随机选数据反而表现集中,挑战传统认知。
神经网络的缩放定律通过数据量评估数据价值,而Vendi Score则利用量子熵衡量。本文证明了常见的缩放定律目标和Vendi Score均为次模函数,并揭示Vendi Score是更广泛的矩阵谱函数类的特例,该类包含确定性点过程(DPP)等多种目标。我们还引入弱矩阵单调函数,推导出一类实用的弱次模矩阵谱函数。提出基于稳态方程的更新机制,避免重复特征分解,使维度为m的嵌入在贪婪优化中边际收益评估效率提升O(m)倍,实测平均加速约35,000倍,使ImageNet-1K规模数据上的Vendi Score直接优化成为可能。在此基础上,我们在固定大小、类别均衡及固定训练预算三种设定下,比较了Vendi Score、DPP、设施选址及三种新型矩阵谱变体对测试性能的预测能力。结果显示,设施选址表现最佳。直接优化还发现,尽管Vendi Score在中等分数范围具有预测性,但高分时反而无法反映下游性能。此外,随机选取的固定大小子集,无论是否类别均衡,其评分与性能均高度集中。最后,我们指出数据价值不仅取决于规模、类别平衡与预算,即使控制这些因素,性能仍呈平滑变化,无明显断点。
原文摘要 · Abstract (English)
Neural scaling laws appraise data through dataset size, while the Vendi Score uses quantum entropy to measure dataset value. We show both that common neural-scaling-law objectives and the Vendi Score are submodular. We further show that the Vendi Score is a special case of a broader class of submodular objectives that we call matrix spectral functions. This also includes determinantal (DPP) objectives, as well as many others. We also introduce weakly matrix monotone functions and show how they lead to weakly submodular matrix spectral functions, yielding a broad family of practical objectives for data appraisal. We develop secular-equation-based updates that avoid repeated eigendecompositions during greedy optimization, reducing marginal-gain evaluation for $m$-dimensional embeddings by an $O(m)$ factor relative to oracle queries. This yields an average empirical speedup of about 35,000x, making direct optimization of the Vendi Score feasible on ImageNet-1K-scale datasets. Thus enabled, we compare how well several objectives predict the value of training subsets for held-out test performance under fixed-size, class-balanced, and fixed training-budget regimes, including the Vendi Score, DPPs, facility location, and three new matrix spectral variants. Across multiple datasets, facility location performs the best. Direct optimization also reveals that, while the Vendi Score is predictive over moderate score ranges, pushing the objective to higher values can make it a poor downstream performance proxy. We also find that uniformly at random fixed-size subsets, both unconstrained and class-balanced, are remarkably concentrated in both appraisal scores and held-out performance. Finally, we show that size, class balance, and training budget do not alone determine data value: even when controlling for these factors, performance ranges smoothly from good to bad.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。