测试大模型当裁判时,发现它对研究型AI的判断准确率不足55%。
Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents?

- 设计干预实验,细粒度检测AI研究代理的错误类型。
- 大模型裁判在证据验证上表现最差,整体准确率低于55%。
- 适合开发可信评估系统的研究人员参考。
深度研究代理正日益自动化复杂的信息检索任务,通过多步推理、工具使用和综合生成基于证据的报告。其重要性上升亟需可扩展且可靠的评估方法,使大语言模型作为裁判成为评估事实准确性、证据使用和推理质量的监督范式。然而,这些裁判在深度研究代理中的可靠性尚不明确,形成一个关键的元评估问题:在部署大模型裁判监督研究代理前,必须先评估裁判自身。现有元评估存在两方面不足:(1)依赖粗粒度、主观的人类偏好一致性;(2)聚焦指令遵循或可验证任务,未涵盖开放式代理执行。为此,我们提出REFLECT(REliable Fine-grained LLM judge Evaluation via Controlled inTervention),一个面向细粒度失败检测的元评估基准。REFLECT定义了过程与结果层面的详细失败模式分类,通过在高质量筛选的代理执行轨迹上进行受控、局部干预,生成可验证、全面且细粒度的评估实例。实验表明,当前大模型裁判仍不可靠:即使最优模型,在推理、工具使用和报告质量失败上的总体准确率也低于55%,尤其在证据验证上表现不佳。我们的分类体系与发现揭示了裁判系统的系统性局限,暴露了成本与可靠性间的权衡,并为构建更可靠的深度研究代理评估流水线提供可操作建议。
原文摘要 · Abstract (English)
Deep research agents increasingly automate complex information-seeking tasks, producing evidence-grounded reports via multi-step reasoning, tool use, and synthesis. Their growing role demands scalable, reliable evaluation, positioning LLM-as-judge as a supervision paradigm for assessing factual accuracy, evidence use, and reasoning quality. Yet the reliability of these judges for deep research agents remains poorly understood, posing a critical meta-evaluation problem: before deploying LLM judges to supervise research agents, we must first evaluate the judges themselves. Existing meta-evaluations fall short in two ways: (1) reliance on coarse, subjective human-preference agreement; (2) focus on instruction-following or verifiable tasks, leaving open-ended agent executions unexplored. To address these gaps, we introduce REFLECT (REliable Fine-grained LLM judge Evaluation via Controlled inTervention), a meta-evaluation benchmark targeting fine-grained failure detection in agentic environments. REFLECT defines a detailed taxonomy of process- and outcome-level failure modes, instantiated by performing controlled and localized interventions on quality-screened agent execution traces. This yields verifiable, comprehensive, and fine-grained instances for validating the judge models. Our experiments show that current LLM judges remain unreliable: even the best-performing models achieve overall accuracies below 55% across reasoning, tool-use, and report-quality failures, with especially poor performance on evidence verification. Together, our taxonomy and findings expose systematic judge limitations, reveal tradeoffs in cost and reliability, and offer actionable guidance for building more reliable evaluation pipelines for deep research agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。