arXiv:2605.15416cs.LGcs.AI2026-05中稿 · ICML被引 1

提升大模型判断可靠性,通过自适应置信度排序实现人机判断一致

Margin-Adaptive Confidence Ranking for Reliable LLM Judgement

论文配图:Margin-Adaptive Confidence Ranking for Reliable LLM Judgement
图 1 · 摘自论文原文
  • 构建基于边界排序的置信度估计器,模拟标注者多样性
  • 在多数据集上显著提升判断一致性成功率,置信度与分歧风险更匹配
  • 适合需要高可靠性的评测场景,如自动化评审系统

Jung等(2025)提出一种假设检验框架,旨在保证大语言模型(LLM)与人类判断的一致性,其前提是模型置信度与人类分歧风险呈单调关系。然而实践中该假设常被违反,且置信度估计器的泛化行为未被明确定义。本文通过学习专用置信度估计器,替代启发式信号,利用模拟标注者多样性和基于边距的排序方法,显式建模模型区分人类一致与分歧案例的能力。我们进一步推导了该估计器的泛化保证,揭示了依赖边距的权衡关系,指导自适应训练流程设计。在固定序列测试中集成该估计器后,排名准确率提升,置信度与分歧风险的单调性增强,在多个数据集和裁判模型下,达成目标一致水平的成功率更高。

原文摘要 · Abstract (English)

Jung et al. (2025) introduce a hypothesis testing framework for guaranteeing agreement between large language models (LLMs) and human judgments, relying on the assumption that the model's estimated confidence is monotonic with respect to human-disagreement risk. In practice, however, this assumption may be violated, and the generalization behavior of the confidence estimator is not explicitly analyzed. We mitigate these issues by learning a dedicated confidence estimator instead of relying on heuristic confidence signals. Our approach leverages simulated annotator diversity and a margin-based ranking formulation to explicitly model how confidently an LLM distinguishes between human-agreement and human-disagreement cases. We further derive generalization guarantees for this estimator, revealing a margin-dependent trade-off that informs the design of an adaptive estimator training procedure. When integrated into fixed-sequence testing, the learned confidence estimator yields improved ranking accuracy and empirically strengthens the monotonic relationship between confidence and disagreement risk, leading to higher success rates in satisfying target agreement levels across multiple datasets and judge models.

大模型评估置信度估计人机对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。