arXiv:2601.13387cs.CLcs.LG2026-01ACL被引 6

用时间逻辑捕捉LLM推理中的信心变化,提升判断准确性。

Confidence over Time: Confidence Calibration with Temporal Logic for Large Language Model Reasoning

  • 用时序逻辑分析每步推理的信心演变模式。
  • 在多个任务上,信心校准度优于现有方法。
  • 适合需要可信推理过程的数学与科学问答场景。

大型语言模型(LLMs)越来越多地依赖长序列、多步骤推理来解决数学问题和科学问答等复杂任务。尽管性能强劲,现有信心估计方法通常将整个推理过程简化为单一标量分数,忽略了信心随生成过程的变化。因此,这些方法容易受响应长度或表达冗余等表面因素影响,难以区分正确推理与自信错误。本文提出使用信号时序逻辑(STL)刻画逐步信心信号。通过判别式STL挖掘程序,发现能区分正确与错误回答信心信号的时间公式。分析表明,这些STL模式在不同任务间具有泛化能力,而数值参数对具体问题敏感。基于此,我们设计了一种信心估计方法,利用超网络为STL模块注入参数。在多个推理任务上的实验表明,该方法获得的信心评分比基线更校准。

原文摘要 · Abstract (English)

Large Language Models (LLMs) increasingly rely on long-form, multi-step reasoning to solve complex tasks such as mathematical problem solving and scientific question answering. Despite strong performance, existing confidence estimation methods typically reduce an entire reasoning process to a single scalar score, ignoring how confidence evolves throughout the generation. As a result, these methods are often sensitive to superficial factors such as response length or verbosity, and struggle to distinguish correct reasoning from confidently stated errors. We propose to characterize the stepwise confidence signal using Signal Temporal Logic (STL). Using a discriminative STL mining procedure, we discover temporal formulas that distinguish confidence signals of correct and incorrect responses. Our analysis found that the STL patterns generalize across tasks, and numeric parameters exhibit sensitivity to individual questions. Based on these insights, we develop a confidence estimation approach that informs STL blocks with parameter hypernetworks. Experiments on multiple reasoning tasks show our confidence scores are more calibrated than the baselines.

LLM推理信心校准时序逻辑

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。