让大模型评判更可信:用相似样本校准评分,防止错误通过。
DA-RAC: Distance-Aware Calibration of LLM Judges for Trustworthy AI Auditing
- 基于语义结构相似性检索参考样例,按距离加权
- 降低虚假通过率,提升评分校准度,优于零样本和固定参考
- 适合高风险生成内容的自动化审核,强调可解释性
生成式AI系统日益产出真实世界成果,但其效果常依赖无上下文的大模型评分。此类评分易受无关参考示例干扰,产生虚假信心,导致低质或有害输出通过评估。我们研究此问题为上下文诱发的校准偏差,提出DA-RAC——一种距离感知的参考锚定校准方法。DA-RAC为每个评估场景检索语义与结构相似的已标注锚点,按距离加权,并将邻域难度作为校准与分诊信号。在多轮大模型评分基准测试中,相比零样本、思维链及静态锚点基线,该方法显著提升校准度并降低误通过风险。机制分析表明,评分随锚点距离系统性变化,而静态参考可能诱导误导性决策边界。因此,大模型评估不仅需更好模型,还需可校准、可审计的参考选择,尤其在高影响生成内容的自动化评估中。评判应基于相关、可检视、可争议的解释性样本。
原文摘要 · Abstract (English)
Generative AI systems are increasingly producing real-world artifacts, however their efficacy and validity are often evaluated via context-free LLM-scoring. These judges can be miscalibrated by irrelevant in-context reference examples, creating false confidence and allowing low-quality or harmful outputs to pass evaluation. We study this failure mode as context-induced miscalibration and introduce DA-RAC, a distance-aware reference-anchored calibration method for LLM judges. DA-RAC retrieves semantically and structurally similar labeled anchors for each judgement scenario, weights them by distance, and exposes neighborhood difficulty as a calibration and triage signal. On multi-run LLM-judge evaluation benchmarks, it improves calibration and reduces false-pass risk relative to zero-shot, chain-of-thought evaluation, and static-anchor baselines. Mechanistic analysis shows that judge scores vary systematically with anchor distance, while static references can induce misleading decision boundaries. Thus LLM-judgement requires not only better models, but also calibrated, auditable reference selection, especially when automated evaluation is used to support high-impact AI generated artifacts. Judgments should be grounded in relevant, inspectable, and contestable interpretive artifacts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。