研究大模型搜索的对抗攻击动态,揭示合作与攻击的博弈规律。
Dynamics of Adversarial Attacks on Large Language Model-Based Search Engines
- 将攻击行为建模为无限重复囚徒困境,分析合作可持续条件。
- 发现降低攻击成功率可能反而激励攻击,防御措施需谨慎设计。
- 适用于关注大模型安全、生态设计的系统研究人员。
大型语言模型(LLM)驱动的搜索系统正重塑信息检索格局,但易受对抗攻击影响,尤其是排名操纵攻击——攻击者通过伪造网页内容操控LLM排序,从而获得不正当竞争优势。本文研究此类攻击的动态机制,将其建模为无限重复囚徒困境,多个参与者在是否合作或攻击间策略抉择。分析表明,合作能否持续取决于攻击成本、贴现率、攻击成功率及触发策略等关键因素。识别出系统动态中的临界点:当参与者更具前瞻性时,合作更易维持。但从防御角度看,单纯降低攻击成功率可能产生悖论效应,反而激励攻击;限制攻击成功率上限的防御手段在某些情境下可能失效。这些发现凸显了保护LLM系统复杂性。本工作提供理论基础与实践洞见,强调自适应安全策略与审慎的生态系统设计的重要性。
原文摘要 · Abstract (English)
The increasing integration of Large Language Model (LLM) based search engines has transformed the landscape of information retrieval. However, these systems are vulnerable to adversarial attacks, especially ranking manipulation attacks, where attackers craft webpage content to manipulate the LLM's ranking and promote specific content, gaining an unfair advantage over competitors. In this paper, we study the dynamics of ranking manipulation attacks. We frame this problem as an Infinitely Repeated Prisoners' Dilemma, where multiple players strategically decide whether to cooperate or attack. We analyze the conditions under which cooperation can be sustained, identifying key factors such as attack costs, discount rates, attack success rates, and trigger strategies that influence player behavior. We identify tipping points in the system dynamics, demonstrating that cooperation is more likely to be sustained when players are forward-looking. However, from a defense perspective, we find that simply reducing attack success probabilities can, paradoxically, incentivize attacks under certain conditions. Furthermore, defensive measures to cap the upper bound of attack success rates may prove futile in some scenarios. These insights highlight the complexity of securing LLM-based systems. Our work provides a theoretical foundation and practical insights for understanding and mitigating their vulnerabilities, while emphasizing the importance of adaptive security strategies and thoughtful ecosystem design.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。