arXiv:2508.05670cs.CRcs.AI2025-08被引 3

测试大模型在攻防博弈中的行为表现,发现语言和性格影响结果

Can LLMs effectively provide game-theoretic-based scenarios for cybersecurity?

  • 构建可复现的博弈型大模型代理框架,测试零和与囚徒困境场景
  • 四款主流大模型在五种语言下均出现收益偏差,受性格与重复次数影响
  • 揭示语言差异对安全应用的潜在风险,建议选型时评估跨语言稳定性

博弈论长期作为网络安全中预测和设计攻防策略互动的基础工具。本文研究经典博弈框架能否有效捕捉由大语言模型驱动的实体行为。基于可复现的博弈型大模型代理框架,我们测试了两类典型场景:一次性零和博弈与动态囚徒困境。实验涵盖四种先进大语言模型,覆盖英语、法语、阿拉伯语、越南语及中文五种自然语言,以评估语言敏感性。结果显示,最终收益受代理个性特征或对重复轮次认知的影响。此外,意外发现最终收益对语言选择高度敏感,警示在不同国家部署大模型时需谨慎,呼吁深入研究其行为差异。我们还引入量化指标评估代理内部一致性与跨语言稳定性,为安全应用中模型选型与优化提供依据。

原文摘要 · Abstract (English)

Game theory has long served as a foundational tool in cybersecurity to test, predict, and design strategic interactions between attackers and defenders. The recent advent of Large Language Models (LLMs) offers new tools and challenges for the security of computer systems; In this work, we investigate whether classical game-theoretic frameworks can effectively capture the behaviours of LLM-driven actors and bots. Using a reproducible framework for game-theoretic LLM agents, we investigate two canonical scenarios -- the one-shot zero-sum game and the dynamic Prisoner's Dilemma -- and we test whether LLMs converge to expected outcomes or exhibit deviations due to embedded biases. Our experiments involve four state-of-the-art LLMs and span five natural languages, English, French, Arabic, Vietnamese, and Mandarin Chinese, to assess linguistic sensitivity. For both games, we observe that the final payoffs are influenced by agents characteristics such as personality traits or knowledge of repeated rounds. Moreover, we uncover an unexpected sensitivity of the final payoffs to the choice of languages, which should warn against indiscriminate application of LLMs in cybersecurity applications and call for in-depth studies, as LLMs may behave differently when deployed in different countries. We also employ quantitative metrics to evaluate the internal consistency and cross-language stability of LLM agents, to help guide the selection of the most stable LLMs and optimising models for secure applications.

博弈论大模型安全多语言攻防对抗

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。