无需真实数据,用统计量体积评估数据集可靠性。
Data Reliability Scoring
- 基于观测数据与实验结果的向量体积构建评分指标
- 在多种噪声模型下准确反映数据质量差异
- 对实验设计不敏感,适合跨场景数据可信度评估
如何在无真实标签的情况下评估数据集的可靠性?我们提出了针对可能来自策略性来源的数据集的可靠性评分问题。真实数据不可见,但可观测到依赖于这些数据的未知统计实验的结果。为基准评估,我们定义了基于真实情况的排序,以衡量报告数据偏离真实值的程度。我们提出格拉姆行列式评分,通过描述观测数据与实验结果经验分布的向量所张成的体积来度量可靠性。该评分保持了多种基于真实情况的可靠性排序,并且唯一地(仅允许缩放)在任意实验下给出相同的可靠性排名,这一性质称为实验无关性。在合成噪声模型、CIFAR-10嵌入和真实就业数据上的实验表明,格拉姆行列式评分能有效捕捉不同观测过程下的数据质量。
原文摘要 · Abstract (English)
How can we assess the reliability of a dataset without access to ground truth? We introduce the problem of reliability scoring for datasets collected from potentially strategic sources. The true data are unobserved, but we see outcomes of an unknown statistical experiment that depends on them. To benchmark reliability, we define ground-truth-based orderings that capture how much reported data deviate from the truth. We then propose the Gram determinant score, which measures the volume spanned by vectors describing the empirical distribution of the observed data and experiment outcomes. We show that this score preserves several ground-truth-based reliability orderings and, uniquely up to scaling, yields the same reliability ranking of datasets regardless of the experimen -- a property we term experiment agnosticism. Experiments on synthetic noise models, CIFAR-10 embeddings, and real employment data demonstrate that the Gram determinant score effectively captures data quality across diverse observation processes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。