研究上下文如何影响大模型决策,发现相同逻辑下不同表述会导致截然不同的结果。
Framing the Game: How Context Shapes LLM Decision-Making
- 通过程序化生成多种情境,系统测试大模型在相同规则下的决策差异。
- 发现大模型响应受上下文描述影响显著,同一问题不同说法结果相差可达30%以上。
- 提示开发者部署时需考虑语境设计,避免误导性表述导致错误判断。
大型语言模型(LLMs)正被广泛应用于多样化的场景中支持决策。尽管现有评估方法能有效探测模型的潜在能力,但常忽视上下文框架对感知理性决策的影响。本研究提出一种新颖的评估框架,通过系统性地改变评估实例的关键特征,并程序化生成情景片段,构建高度多样的决策场景。通过对相同游戏结构下不同上下文中的决策模式进行分析,我们发现大模型的响应存在显著的上下文依赖性。研究结果表明,这种可变性虽具有较高可预测性,却对框架效应极为敏感。该发现强调了在实际应用中需要采用动态、上下文感知的评估方法。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are increasingly deployed across diverse contexts to support decision-making. While existing evaluations effectively probe latent model capabilities, they often overlook the impact of context framing on perceived rational decision-making. In this study, we introduce a novel evaluation framework that systematically varies evaluation instances across key features and procedurally generates vignettes to create highly varied scenarios. By analyzing decision-making patterns across different contexts with the same underlying game structure, we uncover significant contextual variability in LLM responses. Our findings demonstrate that this variability is largely predictable yet highly sensitive to framing effects. Our results underscore the need for dynamic, context-aware evaluation methodologies for real-world deployments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。