用人类仲裁重评基准数据,发现大模型幻觉检测标准可大幅提升。
Do Benchmarks Underestimate LLM Performance? Evaluating Hallucination Detection With LLM-First Human-Adjudicated Assessment
- 通过双文化人类仲裁重评争议样本,修正原标注偏差。
- 模型准确率提升最高达8.51%,三者一致率提高超7%。
- 当模型提供理由时,人类更倾向采纳模型判断,适合高风险场景使用。
幻觉仍是大语言模型在上下文相关任务(如RAG和代理系统)中的主要挑战。本研究聚焦摘要任务中的上下文幻觉检测,对比Gemini 2.5 Flash与GPT-5 Mini在QAGS-C和SummEval数据集上的原始标注与基于原因和片段的预测。针对人类标注与模型判断间的系统性差异,对所有争议样本进行双文化人类仲裁。重评后,三者一致性在QAGS-C和SummEval上分别提升6.38%和7.62%。模型准确率同步提升:GPT在两数据集上分别提升4.25%和2.34%,Gemini分别提升8.51%和3.80%。仲裁员普遍支持有明确推理依据的模型判断,整体人类仲裁一致性为83%–87%。结果表明,对于模糊任务,单次标注不足,模型辅助重评能生成更可靠的评估基准。
原文摘要 · Abstract (English)
Hallucination remains a persistent challenge in Large Language Models (LLMs), particularly in context-grounded settings such as RAG and agentic AI systems. This study focuses on contextual hallucination detection in summarization tasks. We analyze the QAGS-C and SummEval datasets by comparing original benchmark annotations with reason and span-based predictions from Gemini 2.5 Flash and GPT-5 Mini. To address systematic divergences between human labels and LLM judgments, we re-evaluated all conflicted samples through a human adjudication process involving 2 cross-cultural adjudicators. Following this re-evaluation, triple agreement (between human, GPT, and Gemini) increased by 6.38% for QAGS-C and 7.62% for SummEval. Similarly, model accuracy improved, with GPT increasing by 4.25% on QAGS-C and 2.34% on SummEval, while Gemini showed gains of 8.51% and 3.80%, respectively. Notably, adjudicators frequently sided with the models' judgments over original human annotations when LLMs provided explicit reasoning. Overall human adjudicator agreement ranged between 83% and 87%. These findings suggest that for ambiguity-prone tasks, single-pass annotations may be insufficient, and model-assisted re-evaluation yields more reliable benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。