arXiv:2606.29033cs.IR2026-06

让人类专注找关键信息,AI负责匹配,提升大模型评估的可信度。

Human-in-the-Loop Nugget Annotation for Accountable LLM-as-a-Judge Evaluations

论文配图:Human-in-the-Loop Nugget Annotation for Accountable LLM-as-a-Judge Evaluations
图 1 · 摘自论文原文
  • 人类识别输出中的关键信息点(nuggets),AI完成大规模匹配
  • 相比传统方法,减少专家依赖,避免机械认同或认知过载
  • 适合需要可追溯、高可信度评估的AI系统评测场景

可靠评估AI或代理系统的输出需要人类判断,但如何融入人类直接影响结果是真实质量信号还是昂贵形式主义。现有方法要么无意中固化专家偏好(导致机械认同),要么让专家在高认知负荷任务中孤立无援。本文提出一种标注工具原型,采用新分工:人类识别输出中重要的信息片段(nuggets),而大模型负责将这些片段与系统输出进行大规模匹配。该设计发挥双方优势,同时保持真实的人类监督。文中描述了人机协作流程、关键设计决策,并展示了生成的nugget库如何用于自动化评判。

原文摘要 · Abstract (English)

Evaluating AI/Agentic system outputs reliably requires human judgment, but how one incorporates the human determines whether one gets a real quality signal or expensive theater. The common approaches either accidentally anchor human experts (leading to rubber-stamping) or leave them unsupported in cognitively demanding labeling tasks. We present a prototype of an annotation tool that implements a different division of labor: humans identify what information matters (nuggets), while LLMs handle high-volume matching of nuggets to system outputs. This plays to each party's strengths while maintaining genuine human oversight. We describe the Human-AI workflow, key design decisions, and how resulting nugget banks are used with automated judges.

大模型评估人机协同可问责性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。