arXiv:2606.29113cs.GTcs.AI2026-06

用博弈论建模大模型沟通中的欺骗与认知盲区,设计反欺诈机制。

LLM Semantic Signaling Game and Mechanism Design: Systematic Blindness, Awareness Shaping, and Mindset Dynamics

论文配图:LLM Semantic Signaling Game and Mechanism Design: Systematic Blindness, Awareness Shaping, and Mindset Dynamics
图 1 · 摘自论文原文
  • 构建语言信号博弈框架,模拟发送方控制语义、接收方感知反馈的策略互动。
  • 发现接收方认知类型影响判断,重塑其认知可降低钓鱼攻击成功率。
  • 适合研究人机交互安全、大模型通信机制的设计者与安全研究人员。

大型语言模型(LLMs)越来越多地通过自然语言中介战略互动,使语义控制成为沟通与欺骗的关键。本文构建了一个语义信号博弈:发送方选择语义控制,LLM生成随机消息,接收方使用依赖认知状态的评分机制评估消息。接收方的认知被建模为决定其感知和使用哪些语言特征进行推理的类型,从而形式化了系统性盲视。该框架连接了基于提示的控制、统计检测与博弈论均衡分析。对聚合消息得分的高斯近似支持似然比决策规则,而完美贝叶斯纳什均衡刻画了策略行为。论文进一步提出机制设计方法,可重塑接收方认知、惩罚欺骗性语义控制,并调整接收方群体以诱导良性分池均衡。数值实验验证了高斯近似,量化了认知排序效应,在适应性对手下分析了心态动态,并展示了认知塑造与护栏成本如何减少成功钓鱼攻击。所提框架为分析代理型AI系统中战略性语言交互提供了原则性基础,也为设计鲁棒且安全的人机通信提供了新工具。

原文摘要 · Abstract (English)

Large language models (LLMs) increasingly mediate strategic interactions through natural language, making semantic control a critical element of communication and deception. This paper develops a semantic signaling game in which a sender selects a semantic control, an LLM generates a stochastic message, and a receiver evaluates the message using an awareness-dependent scoring mechanism. Receiver awareness is modeled as a type that determines which linguistic features are perceived and used for inference, providing a formal model of systematic blindness. The framework connects prompt-based control, statistical detection, and game-theoretic equilibrium analysis. Gaussian approximations of aggregate message scores enable likelihood-ratio decision rules, while Perfect Bayesian Nash equilibria characterize strategic behavior. The paper further develops mechanism-design approaches that reshape receiver awareness, penalize deceptive semantic controls, and modify receiver populations to induce benign pooling equilibria. Numerical experiments validate the Gaussian approximation, quantify awareness-ordering effects, analyze mindset dynamics under adaptive adversaries, and demonstrate how awareness shaping and guardrail costs reduce successful phishing attacks. The proposed framework provides a principled foundation for analyzing strategic language-mediated interactions in agentic AI systems and offers new tools for the design of robust and secure human-AI communication.

大模型安全博弈论认知建模语言控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。