arXiv:2503.17523cs.CLcs.AI2025-03被引 36

教大模型像贝叶斯一样思考,让它更会根据行为推测用户偏好。

Bayesian Teaching Enables Probabilistic Reasoning in Large Language Models

  • 用贝叶斯模型生成数据,让大模型模仿其推理过程。
  • 改进后的大模型在多轮交互中信念更新准确率显著提升。
  • 适合需要个性化推理的智能助手场景。

大型语言模型(LLMs)越来越多地被用作与用户和世界交互的智能体。为成功完成任务,这些模型必须构建对世界的表征,并形成关于它们的概率性信念。例如,在提供个性化推荐时,模型需从用户的多轮行为中推断其偏好。贝叶斯推理框架为智能体在获取新信息时如何更新信念提供了最优标准。我们首先发现,现有大模型远未达到这一标准。随后,通过训练大模型模仿规范的贝叶斯模型预测,其信念更新能力得到显著提升,且该能力可泛化至新任务。结果表明,大模型能通过示例学习推理技能,并推广到新领域。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly used as agents that interact with users and with the world. To do so successfully, LLMs must construct representations of the world and form probabilistic beliefs about them. To provide personalized recommendations, for example, the LLM needs to infer a user's preferences from their behavior over multiple interactions. The Bayesian inference framework lays out the optimal way for an agent to update its beliefs as it receives new information. We first show that LLMs fall far short of the standard defined by the Bayesian framework. We then show that by teaching LLMs to mimic the predictions of the normative Bayesian model, we can dramatically improve their ability to update their beliefs; this ability generalizes to new tasks. We conclude that LLMs can effectively learn reasoning skills from examples and generalize those skills to new domains.

贝叶斯推理大模型信念更新

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。