研究在人类标注不可靠时如何评估模型,发现现有指标会掩盖问题。
"All that Glitters": Approaches to Evaluations with Unreliable Model and Human Annotations
- 用六维指标评估标注质量,识别出模型与人类标注中的偏差
- 编码器模型看似表现超人,但严格评估后暴露虚假相关与种族偏见
- 在人机协作中,模型可能降低人类标注可靠性,需谨慎使用
人工标注常含误差,这些误差可能隐藏在常规标注质量指标之下,干扰模型评估中的准确性、偏见、公平性与实用性判断。本研究针对课堂教学质量评估这一高成本、依赖人力的任务,分析人类标注、GPT模型评分及变换器编码器模型的标注,采用六维评估框架(一致性、置信度、有效性、偏见、公平性、帮助性)评估两类大语言模型(编码器与GPT解码器)的表现。结果显示,在标准指标下,编码器模型达到甚至超过人类水平;但采用更严格评估方法后,揭示出模型与人类均存在虚假相关性和非随机的种族偏见。研究进一步探讨了在人机协同场景中,若引入模型标注,可能加剧人类标注的方差,降低其可靠性。研究识别出部分模型在当前数据泛化范围内,可提升昂贵的人类教学评价质量。
原文摘要 · Abstract (English)
"Gold" and "ground truth" human-mediated labels have error. The effects of this error can escape commonly reported metrics of label quality or obscure questions of accuracy, bias, fairness, and usefulness during model evaluation. This study demonstrates methods for answering such questions even in the context of very low reliabilities from expert humans. We analyze human labels, GPT model ratings, and transformer encoder model annotations describing the quality of classroom teaching, an important, expensive, and currently only human task. We answer the question of whether such a task can be automated using two Large Language Model (LLM) architecture families--encoders and GPT decoders, using novel approaches to evaluating label quality across six dimensions: Concordance, Confidence, Validity, Bias, Fairness, and Helpfulness. First, we demonstrate that using standard metrics in the presence of poor labels can mask both label and model quality: the encoder family of models achieve state-of-the-art, even "super-human", results across all classroom annotation tasks. But not all these positive results remain after using more rigorous evaluation measures which reveal spurious correlations and nonrandom racial biases across models and humans. This study then expands these methods to estimate how model use would change to human label quality if models were used in a human-in-the-loop context, finding that the variance captured in GPT model labels would worsen reliabilities for humans influenced by these models. We identify areas where some LLMs, within the generalizability of the current data, could improve the quality of expensive human ratings of classroom instruction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。