忽略答案不确定性会误判非专家与专家表现相似。
The Illusion of AI Expertise Under Uncertainty: Navigating Elusive Ground Truth via a Probabilistic Paradigm
- 提出概率范式,量化专家评分受答案不确定性影响的程度。
- 在6个数据集、9个模型上发现:高不确定性下专家与随机标注者表现无显著差异。
- 建议按答案置信度分层评估,提升性能对比可靠性。
评估AI系统能力时,通常忽略专家标注中固有的不确定性。这种模糊性不仅存在于人类偏好,也广泛存在于医疗等安全关键领域。本文提出概率范式,理论上表明:只有在地面真值高度确定时,专家才能获得高分;而在地面真值变异大的数据集中,非专家与专家的表现可能无明显差异。这一现象同样影响模型比较,不确定性掩盖了劣质模型与优质模型之间的差距。因此,忽视地面真值的不确定性可能导致错误结论:非专家表现与专家相当。基于该范式,我们引入期望准确率与期望F1,评估在6个数据集和9个模型上,不同地面真值变异性下专家与系统的表现。结果表明,在专家表现较低时,按地面真值答案概率分层评估至关重要。分层后,高确定性区间内的性能比较更加可靠,有效缓解了不确定性这一关键混杂因素的影响。
原文摘要 · Abstract (English)
Benchmarking the capabilities of AI systems, including Large Language Models (LLMs) and Vision Models, typically ignores the impact of uncertainty in the underlying ground truth answers from experts. This ambiguity is not just limited to human preferences, but is also consequential even in safety critical domains such as medicine where uncertainty is pervasive. In this paper, we introduce a probabilistic paradigm to theoretically explain how high certainty in ground truth answers is almost always necessary for even an expert to achieve high scores, whereas in datasets with high variation in ground truth answers there may be little difference between a random labeller and an expert. This characteristic also manifests when comparing models, where uncertainty obfuscates differences between poor and high performing models. Therefore, ignoring uncertainty in ground truth evaluation data can result in the misleading conclusion that a non-expert has similar performance to that of an expert. Using the probabilistic paradigm, we thus bring forth the concepts of expected accuracy and expected F1 and compare the estimated score an expert human or system can achieve given ground truth answer variability across 6 datasets and 9 models. The results lead to the recommendation that stratification by the probability of the ground truth answer becomes critical when expert performance is relatively low. Under stratified evaluation, performance comparison becomes more reliable in high certainty bins, mitigating the effect of the key confounding factor -- uncertainty.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。