用自评排序稳定性衡量大模型推理一致性,发现可靠推理有稳定路径。
Quantifying Consistency in LLM Logical Reasoning via Structural Uncertainty

- 让模型自评多个推理路径的优劣,用排名分布判断一致性。
- 在逻辑与数学任务中,结合结构不确定性可更准识别不可靠答案。
- 适合关注推理可信度的研究者,尤其多步推理解析场景。
大语言模型虽常得出相同答案,但推理路径可能不稳定、矛盾或难以一致排序,尤其在多步演绎推理中尤为明显。现有方法主要依赖输出分散度评估可靠性,却忽略了模型对自身推理候选解的一致性排序能力。本文提出结构不确定性框架,基于模型自生成解的偏好排序稳定性构建。针对同一问题生成多个候选解,让模型进行两两比较偏好,通过Bradley-Terry模型与PageRank聚合为排名分布,并分解为两个熵基成分:跨试验排序不稳定性与单次试验内候选模糊性。在五种LLM和八个基准测试中,结构信号补充了答案分散度信息:在逻辑与数学推理任务中,组合使用提升不可靠实例识别能力;而在事实检索任务中,结构信号趋于均匀,揭示了推理一致性评估失效的临界区间。两个分量与准确率关系不同:内部模糊性与正确性正相关(多条合理路径并存),跨试验不稳定性则负相关,指示推理不可靠。结构不确定性并非通用置信度指标,而是依赖场景的逻辑推理一致性评估器。
原文摘要 · Abstract (English)
Large language models can arrive at the same answer through reasoning paths that are unstable, contradictory, or difficult to rank consistently -- a failure mode especially prevalent in multi-step deductive reasoning. Existing methods assess reliability primarily through output dispersion -- measuring how much sampled answers differ -- but this discards a complementary signal: whether the model can consistently rank competing reasoning candidates. We propose structural uncertainty, a consistency-aware framework derived from the stability of self-preference-induced rankings over sampled reasoning solutions. Given a query, we generate multiple candidate solutions and ask the model to judge pairwise preferences among its own outputs. We aggregate self-preferences into ranking distributions via Bradley-Terry modeling with PageRank, and decompose the signal into two entropy-based components: across-trial ranking instability and within-trial candidate ambiguity. Across five LLMs and eight benchmarks, structural signals provide information complementary to answer dispersion: on logical and mathematical reasoning tasks, the combination improves identification of unreliable instances, while on factual retrieval the structural signal collapses toward uniformity, diagnosing a regime boundary where reasoning-level consistency evaluation is uninformative. The two components relate differently to accuracy: within-trial ambiguity correlates positively with correctness -- consistent with settings where multiple plausible solution paths remain competitive -- while across-trial instability correlates negatively, signaling unreliable reasoning. Structural uncertainty is best understood not as a universal confidence estimator, but as a regime-sensitive evaluator of logical reasoning consistency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。