研究大模型在不同训练阶段的风险偏好,提出有效调控方法。
Risk Profiling and Modulation for LLMs
- 结合行为经济学,构建评估大模型风险偏好的新框架。
- 指令微调模型符合标准效用理论,而强化学习对齐模型偏差更大。
- 后训练是稳定调控模型风险偏好的最有效手段,适合风险敏感场景。
大语言模型(LLMs)越来越多地用于不确定环境下的决策任务,但其风险特征及提示和对齐方法的影响仍不明确。现有研究主要关注人格化提示或多智能体交互,未深入探讨后训练如何影响模型风险行为。本文提出一种新流程,用于诱发、引导和调节大模型的风险偏好,借鉴行为经济学与金融学工具。通过效用理论模型,对比预训练、指令微调和基于强化学习人类反馈(RLHF)对齐的模型,发现指令微调模型的行为符合部分标准效用形式,而预训练和RLHF对齐模型更偏离拟合的效用模型。进一步评估提示工程、上下文学习和后训练等调控策略,结果表明后训练在风险偏好的稳定性和有效性上表现最佳。研究揭示了不同类别与训练阶段大模型的风险特征,并展示后训练如何调节这些特征,为未来行为对齐与风险感知大模型设计奠定基础。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly used for decision-making tasks under uncertainty; however, their risk profiles and how they are influenced by prompting and alignment methods remain underexplored. Existing studies have primarily examined personality prompting or multi-agent interactions, leaving open the question of how post-training influences the risk behavior of LLMs. In this work, we propose a new pipeline for eliciting, steering, and modulating LLMs' risk profiles, drawing on tools from behavioral economics and finance. Using utility-theoretic models, we compare pre-trained, instruction-tuned, and RLHF-aligned LLMs, and find that while instruction-tuned models exhibit behaviors consistent with some standard utility formulations, pre-trained and RLHF-aligned models deviate more from any utility models fitted. We further evaluate modulation strategies, including prompt engineering, in-context learning, and post-training, and show that post-training provides the most stable and effective modulation of risk preference. Our findings provide insights into the risk profiles of different classes and stages of LLMs and demonstrate how post-training modulates these profiles, laying the groundwork for future research on behavioral alignment and risk-aware LLM design.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。