arXiv:2606.19714stat.MLcs.AI2026-06

让AI judge自我修正,提升评估可靠性。

AURA: Adaptive Uncertainty-aware Refinement for LLM-as-a-Judge Auditing

论文配图:AURA: Adaptive Uncertainty-aware Refinement for LLM-as-a-Judge Auditing
图 1 · 摘自论文原文
  • 基于不确定性动态筛选需人工验证的对比项。
  • 在真实数据上使评估一致性提升12.7%。
  • 适合需要高可信度AI评估的场景。

大语言模型(LLM)被广泛用作开放生成任务的评判者,因其相比大规模人工评价更具可扩展性,但其判断仍无法完全替代人类。现有审计流程通常假设可预先获得可靠样本或干净监督信号,如人工标注、启发式过滤或强裁判输出。然而在LLM评估中,这种假设不可靠:初始划分可能继承裁判偏见,而人工验证又常因稀缺难以形成稳定群体。本文提出AURA框架,一种自适应不确定性感知的精炼机制,在有限人工验证下优化成对的LLM作为裁判的决策。AURA通过迭代学习人类一致性信号,传播可靠证据,并优先将不确定的对比项提交人工审查。核心思想是将对裁判的信任视为随证据积累而逐步精炼的隐变量。我们提供了简洁的公式化表达、稳定的精炼流程,并在合成与真实成对的LLM回答数据上进行了全面评估。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly used as judges for open-ended generation, as large-scale human evaluation is often expensive and difficult to scale, yet their preferences remain imperfect proxies for human judgment. Existing auditing pipelines often assume that a reliable subset of examples or clean supervision signals are available beforehand, for example from human annotation, heuristic filtering, or the outputs of strong judges. In LLM evaluation, this assumption is fragile: the initial split may inherit judge bias, while human verification is typically too scarce to define stable groups at scale. We propose AURA, an adaptive uncertainty--aware refinement framework for auditing pairwise LLM--as--a--judge decisions under selected human verification. AURA iteratively learns a human-consistency signal, propagates reliable evidence, and prioritizes uncertain comparisons for human review. The key idea is to treat trust in a judge as a latent quantity that is progressively refined as evidence accumulates. We provide a compact formulation, a stable refinement procedure, and a comprehensive evaluation on both synthetic and real pairwise LLM-answer data.

AI评估大模型自适应

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。