arXiv:2602.08552cs.LGeess.AS2026-02

提出衡量主观评价数据集最高可实现相关性的方法

Rho-Perfect: Correlation Ceiling For Subjective Evaluation Datasets

  • 定义理想预测器与人工评分的相关性为ρ-Perfect
  • 基于异方差噪声模型推导出该相关性的估算值
  • 可区分模型缺陷与数据质量问题,适合评估评测集可靠性

主观评分天然存在噪声,限制了模型与人类评分的相关性,但这一可靠性问题很少被量化。本文提出ρ-Perfect,即理想预测器与人工评分之间的最高可能相关性,并基于异方差噪声场景(主观评分数据中常见)推导其估计值。我们证明ρ-Perfect的平方可估计测试-重测相关性,并以此验证估计的合理性。在语音质量数据集上的实验表明,ρ-Perfect可用于区分模型性能瓶颈与数据本身的质量问题。

原文摘要 · Abstract (English)

Subjective ratings contain inherent noise that limits the model-human correlation, but this reliability issue is rarely quantified. In this paper, we present $ρ$-Perfect, a practical estimation of the highest achievable correlation of a model on subjectively rated datasets. We define $ρ$-Perfect to be the correlation between a perfect predictor and human ratings, and derive an estimate of the value based on heteroscedastic noise scenarios, a common occurrence in subjectively rated datasets. We show that $ρ$-Perfect squared estimates test-retest correlation and use this to validate the estimate. We demonstrate the use of $ρ$-Perfect on a speech quality dataset and show how the measure can distinguish between model limitations and data quality issues.

主观评价相关性分析评测集质量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。