arXiv:2505.11614cs.AIcs.CL2025-05被引 10

用强化学习让大模型学会解释人类决策,既准又可懂。

Using Reinforcement Learning to Train Large Language Models to Explain Human Decisions

  • 用基于结果的奖励机制训练大模型生成决策推理过程
  • 模型预测人类风险选择表现优异,同时生成高质量自然语言解释
  • 适合想理解人类决策机制的研究者或应用开发者

认知建模的核心目标是构建不仅能预测人类行为,还能揭示其背后认知机制的模型。尽管大规模行为数据训练的神经网络模型通常具备强预测能力,却往往难以提供可解释的认知过程分析。本文探索预训练大语言模型(LLMs)作为双重功能认知模型的潜力——既能实现高精度预测,又能以自然语言生成可解释的推理过程。具体地,我们采用基于结果的强化学习方法,引导大模型生成解释人类风险决策的显式推理轨迹。实验表明,该方法在保持对人类决策强量化预测能力的同时,能生成高质量解释性文本。

原文摘要 · Abstract (English)

A central goal of cognitive modeling is to develop models that not only predict human behavior but also provide insight into the underlying cognitive mechanisms. While neural network models trained on large-scale behavioral data often achieve strong predictive performance, they typically fall short in offering interpretable explanations of the cognitive processes they capture. In this work, we explore the potential of pretrained large language models (LLMs) to serve as dual-purpose cognitive models--capable of both accurate prediction and interpretable explanation in natural language. Specifically, we employ reinforcement learning with outcome-based rewards to guide LLMs toward generating explicit reasoning traces for explaining human risky choices. Our findings demonstrate that this approach produces high-quality explanations alongside strong quantitative predictions of human decisions.

大模型认知建模强化学习可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。