arXiv:2509.10739cs.CL2025-09被引 7

首次系统评估大模型在概率推理中的表现,发现大模型有较强推理能力但受符号表达和上下文长度影响。

Reasoning Under Uncertainty: Exploring Probabilistic Reasoning Capabilities of LLMs

  • 设计三类任务测试模型对概率分布的推理能力
  • 大模型在样本生成上表现优异,小模型差距明显
  • 符号表示变化导致性能下降,长上下文时性能跌超60%

尽管大型语言模型在语言理解与生成方面取得广泛应用,但在需要概率推理的任务中仍表现出不明确且不一致的行为。本文首次系统研究了大模型在显式离散概率分布上的推理能力。通过设计三类任务——模式识别、最大似然估计和样本生成,我们考察模型对联合分布或条件分布的响应能力,以探查频率分析、边缘化及生成行为等多方面技能。大规模实证评估表明,大模型与小模型之间存在显著性能差距,大模型展现出更强的推理能力,尤其在样本生成方面表现突出。然而,研究也揭示若干局限:模型对概率表示符号的变化敏感,随着上下文长度增加,性能下降超过60%。这些结果为理解大模型的概率推理能力提供了详细依据,并指明了未来改进的关键方向。

原文摘要 · Abstract (English)

Despite widespread success in language understanding and generation, large language models (LLMs) exhibit unclear and often inconsistent behavior when faced with tasks that require probabilistic reasoning. In this work, we present the first comprehensive study of the reasoning capabilities of LLMs over explicit discrete probability distributions. Given observations from a probability distribution, we evaluate models on three carefully designed tasks, mode identification, maximum likelihood estimation, and sample generation, by prompting them to provide responses to queries about either the joint distribution or its conditionals. These tasks thus probe a range of probabilistic skills, including frequency analysis, marginalization, and generative behavior. Through comprehensive empirical evaluations, we demonstrate that there exists a clear performance gap between smaller and larger models, with the latter demonstrating stronger inference and surprising capabilities in sample generation. Furthermore, our investigations reveal notable limitations, including sensitivity to variations in the notation utilized to represent probabilistic outcomes and performance degradation of over 60% as context length increases. Together, our results provide a detailed understanding of the probabilistic reasoning abilities of LLMs and identify key directions for future improvement.

概率推理大模型生成能力语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。