arXiv:2602.23546cs.CLcs.AI2026-02Transactions of th…

人类与大模型在概率推理上差异显著,揭示了当前模型的局限性。

Humans and LLMs Diverge on Probabilistic Inferences

  • 构建210个手写概率推理题,由25-30人标注可信度
  • 大模型输出分布与人类差异明显,无法模仿人类渐进判断
  • 发现模型共用相似推理路径,提示其思维机制不同

人类推理常基于有限信息进行概率判断,而非严格逻辑推导。尽管大模型在逻辑与数学任务中表现优异,但在开放性、非确定性推理方面仍缺乏研究。本文提出ProbCOPA数据集,包含210个英文概率推理题,每题由25至30名参与者标注推理可信度。结果表明,人类判断呈梯度变化,体现真实概率认知。对比八种顶尖推理大模型,发现其输出分布与人类存在系统性偏差。进一步分析模型推理链,发现其采用一致的评估模式。研究揭示人类与大模型在概率推理上的根本差异,强调需在非确定性场景下评估模型推理能力。

原文摘要 · Abstract (English)

Human reasoning often involves working over limited information to arrive at probabilistic conclusions. In its simplest form, this involves making an inference that is not strictly entailed by a premise, but rather only likely given the premise. While reasoning LLMs have demonstrated strong performance on logical and mathematical tasks, their behavior on such open-ended, non-deterministic inferences remains largely unexplored. We introduce ProbCOPA, a dataset of 210 handcrafted probabilistic inferences in English, each annotated for inference likelihood by 25--30 human participants. We find that human responses are graded and varied, revealing probabilistic judgments of the inferences in our dataset. Comparing these judgments with responses from eight state-of-the-art reasoning LLMs, we show that models consistently fail to produce human-like distributions. Finally, analyzing LLM reasoning chains, we find evidence of a common reasoning pattern used to evaluate such inferences. Our findings reveal persistent differences between humans and LLMs, and underscore the need to evaluate reasoning beyond deterministic settings.

概率推理大模型评测认知差异

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。