数据估值存在偏见与不稳,建议用透明卡片机制提升可信度。
A case for data valuation transparency via DValCards
- 提出数据估值卡片(DValCards)框架,增强估值过程透明度。
- 预处理和采样会显著改变数据估值结果,导致偏差。
- 低估少数群体数据价值,影响公平性,适合关注数据伦理者。
随着以数据为中心的机器学习兴起,多种数据估值方法被提出,用于量化每个数据点对模型性能(如准确率)的贡献。除了技术应用(如数据清洗、数据采购),有人建议在数据市场中,数据买家可用这些方法公平补偿数据提供者。本文通过分析9个表格分类数据集和6种估值方法,揭示:(1)常见且低成本的数据预处理会大幅改变数据估值;(2)基于估值的子采样可能加剧类别不平衡;(3)估值指标可能低估少数群体数据的价值。这些结果表明数据估值本身存在内在偏见与不稳定性,带来技术和伦理问题。因此,我们主张提升数据估值在实际应用中的透明度,并提出新的数据估值卡片(DValCards)框架。推广该框架可减少估值指标滥用(如定价),并建立负责任机器学习系统的信任基础。
原文摘要 · Abstract (English)
Following the rise in popularity of data-centric machine learning (ML), various data valuation methods have been proposed to quantify the contribution of each datapoint to desired ML model performance metrics (e.g., accuracy). Beyond the technical applications of data valuation methods (e.g., data cleaning, data acquisition, etc.), it has been suggested that within the context of data markets, data buyers might utilize such methods to fairly compensate data owners. Here we demonstrate that data valuation metrics are inherently biased and unstable under simple algorithmic design choices, resulting in both technical and ethical implications. By analyzing 9 tabular classification datasets and 6 data valuation methods, we illustrate how (1) common and inexpensive data pre-processing techniques can drastically alter estimated data values; (2) subsampling via data valuation metrics may increase class imbalance; and (3) data valuation metrics may undervalue underrepresented group data. Consequently, we argue in favor of increased transparency associated with data valuation in-the-wild and introduce the novel Data Valuation Cards (DValCards) framework towards this aim. The proliferation of DValCards will reduce misuse of data valuation metrics, including in data pricing, and build trust in responsible ML systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。