arXiv:2510.23409cs.LGcs.AI2025-10被引 1

提出一种轻量级方法,让数据估值在分布外场景下依然可靠。

Eigen-Value: Efficient Domain-Robust Data Valuation via Eigenvalue-Based Approach

  • 基于特征协方差矩阵的特征值比,快速估算域间差异。
  • 在多个真实数据集上实现稳定且鲁棒的数据价值排序。
  • 无需额外训练,可无缝集成到现有估值方法中。

数据估值在以数据为中心的AI时代日益关键,可驱动高效训练流程并实现数据市场的客观定价。现有方法多基于分布内(ID)验证性能变化来评估数据移除影响,但当验证集不含分布外(OOD)数据时,其结果难以泛化。尽管已有面向OOD的方法,但计算开销大,难以部署。本文提出 extit{Eigen-Value}(EV),一种仅需ID数据子集(包括验证阶段)的即插即用框架。EV通过ID数据协方差矩阵的特征值比,提供域差异的新型谱近似;再利用微扰理论估计每条数据对这一差异的边际贡献,显著降低计算成本。随后,EV通过添加一个无额外训练循环的EV项,集成至原有基于ID损失的方法中。实验表明,该方法在真实数据集上实现了更强的OOD鲁棒性与稳定的价值排序,兼具高效性,适用于存在域偏移的大规模场景。

原文摘要 · Abstract (English)

Data valuation has become central in the era of data-centric AI. It drives efficient training pipelines and enables objective pricing in data markets by assigning a numeric value to each data point. Most existing data valuation methods estimate the effect of removing individual data points by evaluating changes in model validation performance under in-distribution (ID) settings, as opposed to out-of-distribution (OOD) scenarios where data follow different patterns. Since ID and OOD data behave differently, data valuation methods based on ID loss often fail to generalize to OOD settings, particularly when the validation set contains no OOD data. Furthermore, although OOD-aware methods exist, they involve heavy computational costs, which hinder practical deployment. To address these challenges, we introduce \emph{Eigen-Value} (EV), a plug-and-play data valuation framework for OOD robustness that uses only an ID data subset, including during validation. EV provides a new spectral approximation of domain discrepancy, which is the gap of loss between ID and OOD using ratios of eigenvalues of ID data's covariance matrix. EV then estimates the marginal contribution of each data point to this discrepancy via perturbation theory, alleviating the computational burden. Subsequently, EV plugs into ID loss-based methods by adding an EV term without any additional training loop. We demonstrate that EV achieves improved OOD robustness and stable value rankings across real-world datasets, while remaining computationally lightweight. These results indicate that EV is practical for large-scale settings with domain shift, offering an efficient path to OOD-robust data valuation.

数据估值域鲁棒性轻量计算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。