用大模型模拟欺骗博弈,揭示人类认知与策略对抗的深层机制。
LLM-Stackelberg Games: Conjectural Reasoning Equilibria and Their Applications to Spearphishing
- 将大模型引入领导者-追随者博弈,通过提示词实现分层推理与策略调整。
- 提出推测性推理均衡,能捕捉对手行为不确定性,适应信息不对称场景。
- 在钓鱼攻击案例中验证有效性,适用于网络安全与虚假信息研究。
我们提出LLM-Stackelberg博弈框架,一种将大语言模型(LLMs)融入领导者与追随者之间序列决策的新型模型。不同于经典假设中的完全信息与理性代理,该框架允许各参与者通过结构化提示进行推理,利用LLMs生成概率性行为,并基于内部认知与信念更新动态调整策略。我们定义了两种均衡概念:推理与行为均衡,使代理的提示式内部推理与其可观测行为保持一致;以及推测性推理均衡,通过对手响应的概率化建模来应对认知不确定性。这些多层次结构能够刻画有限理性、信息不对称和元认知适应。通过钓鱼邮件攻击案例研究,展示了大模型驱动互动的认知丰富性与对抗潜力。结果表明,该框架为网络安全、虚假信息传播及推荐系统等领域的决策建模提供了有力工具。
原文摘要 · Abstract (English)
We introduce the framework of LLM-Stackelberg games, a class of sequential decision-making models that integrate large language models (LLMs) into strategic interactions between a leader and a follower. Departing from classical Stackelberg assumptions of complete information and rational agents, our formulation allows each agent to reason through structured prompts, generate probabilistic behaviors via LLMs, and adapt their strategies through internal cognition and belief updates. We define two equilibrium concepts: reasoning and behavioral equilibrium, which aligns an agent's internal prompt-based reasoning with observable behavior, and conjectural reasoning equilibrium, which accounts for epistemic uncertainty through parameterized models over an opponent's response. These layered constructs capture bounded rationality, asymmetric information, and meta-cognitive adaptation. We illustrate the framework through a spearphishing case study, where a sender and a recipient engage in a deception game using structured reasoning prompts. This example highlights the cognitive richness and adversarial potential of LLM-mediated interactions. Our results show that LLM-Stackelberg games provide a powerful paradigm for modeling decision-making in domains such as cybersecurity, misinformation, and recommendation systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。