arXiv:2508.19288cs.CRcs.AI2025-08
攻击者可诱导游戏AI角色泄露隐藏剧情秘密
Tricking LLM-Based NPCs into Spilling Secrets
- 通过对抗性提示注入操控游戏AI对话
- 成功让多个基于LLM的非玩家角色泄露预设秘密
- 揭示了生成式AI在游戏中的安全风险,适合安全研究者关注
大型语言模型(LLMs)正被广泛用于生成游戏非玩家角色(NPC)的动态对话。然而,其集成带来了新的安全挑战。本研究探讨了对抗性提示注入是否会导致基于LLM的NPC暴露原本应保密的隐藏背景信息。实验表明,攻击者可通过精心设计的恶意提示,诱导多个典型游戏场景中的LLM NPC透露预设的秘密内容,验证了当前LLM驱动角色在安全机制上的脆弱性。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are increasingly used to generate dynamic dialogue for game NPCs. However, their integration raises new security concerns. In this study, we examine whether adversarial prompt injection can cause LLM-based NPCs to reveal hidden background secrets that are meant to remain undisclosed.
游戏AI安全风险提示攻击
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。