arXiv:2607.10251cs.AI2026-07

用德州扑克测试大模型风险决策行为,发现其有稳定且可区分的风险风格。

Behavioural Signatures of Risk-Sensitive Decision-Making in Large Language Models

论文配图:Behavioural Signatures of Risk-Sensitive Decision-Making in Large Language Models
图 1 · 摘自论文原文
  • 通过参与度和主动度量化模型在不确定性中的决策行为。
  • 不同模型呈现从保守到激进的稳定风险谱,且在混合对局中差异更明显。
  • 适合关注AI决策可解释性与风险审计的研究者阅读。

随着大语言模型(LLMs)越来越多地应用于决策支持,理解其在不确定性下的选择是否具有稳定且可解释的行为规律至关重要。人类决策既包含相对持久的风险偏好,也存在情境依赖的调整,但尚不清楚类似行为结构是否存在于基于LLM的决策系统中。本文基于无限制德州扑克构建了一个受控多模型框架,以参与度(Participation)衡量自愿参与不确定机会的程度,以主动性(Proactiveness)衡量前翻牌阶段的风险升级行为。在同质自对战与异质混合模型交互中,前沿LLMs展现出稳定且模型特异的风险特征,形成从保守到激进的连续谱。这些特征在对手构成变化时仍保持相对稳健,而最保守与最激进模型在混合设置下进一步分化。在全局风险压力和个人资源约束下,模型以结构化但异质的方式适应,表现为整体行为收缩、选择性降级或近乎不变的行为。研究结果表明,LLMs不仅在基础风险倾向上存在差异,还对不同风险信号的响应方式及调整灵活性各不相同,为交互场景下的风险敏感决策审计提供了行为基础。代码已公开:https://github.com/XuankunRong/AgentTexasPoker。

原文摘要 · Abstract (English)

As large language models (LLMs) are increasingly used in decision support, it is important to understand whether their choices under uncertainty exhibit stable and interpretable behavioural regularities. Human decision-making combines relatively persistent risk preferences with context-dependent adjustment, yet it remains unclear whether analogous behavioural structure can be observed in LLM-based decision systems. Here we examine this question using a controlled multi-model framework based on no-limit Texas Hold'em, where behaviour is quantified by Participation, measuring voluntary engagement in uncertain opportunities, and Proactiveness, measuring pre-flop risk escalation. Across homogeneous self-play and heterogeneous mixed-model interactions, frontier LLMs exhibit stable, model-specific risk profiles, forming a spectrum from conservative to aggressive decision styles. These profiles remain largely robust under changing opponent composition, while the most conservative and most aggressive models diverge further in mixed settings. Under global risk pressure and personal resource constraint, models adapt in structured but heterogeneous ways, ranging from broad behavioural contraction to selective de-escalation and near-invariant behaviour. These findings suggest that LLMs differ not only in baseline risk disposition, but also in the risk signals they respond to and the flexibility with which they adjust, providing a behavioural basis for auditing risk-sensitive decision-making in interactive settings. Our code is publicly available at: https://github.com/XuankunRong/AgentTexasPoker.

大模型决策行为风险感知博弈模拟

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。