arXiv:2511.07364cs.LGcs.AI2025-11中稿 · NeurIPS被引 4

让大模型分步评估自身推理,提升复杂任务中的错误检测能力

Self-Evaluating LLMs for Multi-Step Tasks: Stepwise Confidence Estimation for Failure Detection

  • 分步评估比整体评估更有效,逐环节判断可靠性
  • 在多步任务中,错误检测的AUC-ROC最高提升15%相对值
  • 适合需要高可靠性的复杂推理场景,如医疗、金融

大语言模型在高风险多步推理任务中的可靠性与故障检测至关重要。现有方法多聚焦单步输出的置信度估计,忽视多步推理的挑战。本文将自评估技术扩展至多步任务,测试整体评分与分步评分两种策略。基于两个多步基准数据集的实验表明,分步评估在错误检测上普遍优于整体评分,AUC-ROC最高提升15%相对值。结果证明,自评估大模型系统能在复杂推理中提供有意义的置信度估计,增强可信度,并为故障检测提供实用框架。

原文摘要 · Abstract (English)

Reliability and failure detection of large language models (LLMs) is critical for their deployment in high-stakes, multi-step reasoning tasks. Prior work explores confidence estimation for self-evaluating LLM-scorer systems, with confidence scorers estimating the likelihood of errors in LLM responses. However, most methods focus on single-step outputs and overlook the challenges of multi-step reasoning. In this work, we extend self-evaluation techniques to multi-step tasks, testing two intuitive approaches: holistic scoring and step-by-step scoring. Using two multi-step benchmark datasets, we show that stepwise evaluation generally outperforms holistic scoring in detecting potential errors, with up to 15% relative increase in AUC-ROC. Our findings demonstrate that self-evaluating LLM systems provide meaningful confidence estimates in complex reasoning, improving their trustworthiness and providing a practical framework for failure detection.

大模型评估多步推理置信度估计故障检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。