arXiv:2608.14950cs.CL2026-08

让大模型评判更可信:用相似样本校准评分,防止错误通过。

DA-RAC: Distance-Aware Calibration of LLM Judges for Trustworthy AI Auditing

  • 基于语义结构相似性检索参考样例,按距离加权
  • 降低虚假通过率,提升评分校准度,优于零样本和固定参考
  • 适合高风险生成内容的自动化审核,强调可解释性

生成式AI系统日益产出真实世界成果,但其效果常依赖无上下文的大模型评分。此类评分易受无关参考示例干扰,产生虚假信心,导致低质或有害输出通过评估。我们研究此问题为上下文诱发的校准偏差,提出DA-RAC——一种距离感知的参考锚定校准方法。DA-RAC为每个评估场景检索语义与结构相似的已标注锚点,按距离加权,并将邻域难度作为校准与分诊信号。在多轮大模型评分基准测试中,相比零样本、思维链及静态锚点基线,该方法显著提升校准度并降低误通过风险。机制分析表明,评分随锚点距离系统性变化,而静态参考可能诱导误导性决策边界。因此,大模型评估不仅需更好模型,还需可校准、可审计的参考选择,尤其在高影响生成内容的自动化评估中。评判应基于相关、可检视、可争议的解释性样本。

原文摘要 · Abstract (English)

Generative AI systems are increasingly producing real-world artifacts, however their efficacy and validity are often evaluated via context-free LLM-scoring. These judges can be miscalibrated by irrelevant in-context reference examples, creating false confidence and allowing low-quality or harmful outputs to pass evaluation. We study this failure mode as context-induced miscalibration and introduce DA-RAC, a distance-aware reference-anchored calibration method for LLM judges. DA-RAC retrieves semantically and structurally similar labeled anchors for each judgement scenario, weights them by distance, and exposes neighborhood difficulty as a calibration and triage signal. On multi-run LLM-judge evaluation benchmarks, it improves calibration and reduces false-pass risk relative to zero-shot, chain-of-thought evaluation, and static-anchor baselines. Mechanistic analysis shows that judge scores vary systematically with anchor distance, while static references can induce misleading decision boundaries. Thus LLM-judgement requires not only better models, but also calibrated, auditable reference selection, especially when automated evaluation is used to support high-impact AI generated artifacts. Judgments should be grounded in relevant, inspectable, and contestable interpretive artifacts.

大模型评估可信审计校准方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。