arXiv:2512.02719cs.CLcs.AI2025-12

发现大模型能自发进行贝叶斯推理,但准确率高不等于鲁棒。

Emergent Bayesian Behaviour and Optimal Cue Combination in LLMs

  • 用心理物理学实验测大模型多模态信号融合策略
  • GPT-5 Mini准确率满分却无法有效整合视觉信息
  • 提出贝叶斯一致性评分,识别隐藏的最优推理行为

大语言模型在显式推理上表现优异,但其隐式计算机制仍不清晰。心理学研究表明人类在感知任务中能近乎最优地处理噪声信号并进行贝叶斯整合。我们探究大模型是否具备类似能力,且无需显式训练即可实现最优多模态融合。基于心理物理学范式,设计贝叶斯基准测试(BayesBench):四个基于文本和图像的量值估计任务(长度、位置、距离、时长),评估九种大模型与人类判断的校准度。通过控制噪声、上下文和提示词的消融实验,测量多模态线索融合的表现、行为模式与效率。除准确率与效率外,引入贝叶斯一致性评分,可在准确率饱和时检测贝叶斯一致的行为变化。结果表明,虽部分优秀模型能以贝叶斯一致方式适应,但准确率不能保证鲁棒性。值得注意的是,GPT-5 Mini 在文本任务中达到完美准确率,却未能高效整合视觉线索。这揭示了能力与策略间的显著分离,说明以准确率为唯一指标的评测可能忽略脆弱的不确定性处理机制。研究揭示了大模型中涌现的原理性不确定性处理,并发现准确率与贝叶斯倾向存在相关性。我们公开发布心理物理学基准与一致性评分工具(https://bayes-bench.github.io),以支持未来多模态架构设计。

原文摘要 · Abstract (English)

Large language models (LLMs) excel at explicit reasoning, but their implicit computational strategies remain underexplored. Decades of psychophysics research show that humans intuitively process and integrate noisy signals using near-optimal Bayesian strategies in perceptual tasks. We ask whether LLMs exhibit similar behaviour and perform optimal multimodal integration without explicit training or instruction. Adopting the psychophysics paradigm, we infer computational principles of LLMs from systematic behavioural studies. We introduce a behavioural benchmark - BayesBench: four magnitude estimation tasks (length, location, distance, and duration) over text and image, inspired by classic psychophysics, and evaluate a diverse set of nine LLMs alongside human judgments for calibration. Through controlled ablations of noise, context, and instruction prompts, we measure performance, behaviour and efficiency in multimodal cue-combination. Beyond accuracy and efficiency metrics, we introduce a Bayesian Consistency Score that detects Bayes-consistent behavioural shifts even when accuracy saturates. Our results show that while capable models often adapt in Bayes-consistent ways, accuracy does not guarantee robustness. Notably, GPT-5 Mini achieves perfect text accuracy but fails to integrate visual cues efficiently. This reveals a critical dissociation between capability and strategy, suggesting accuracy-centric benchmarks may over-index on performance while missing brittle uncertainty handling. These findings reveal emergent principled handling of uncertainty and highlight the correlation between accuracy and Bayesian tendencies. We release our psychophysics benchmark and consistency metric (https://bayes-bench.github.io) as evaluation tools and to inform future multimodal architecture designs.

贝叶斯推理多模态大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。