剖析大模型推理能力中的后门攻击,揭示安全风险与防御方向。
Rethinking Reasoning: A Survey on Reasoning-based Backdoors in LLMs
- 按关联、被动、主动三类划分推理型后门攻击机制。
- 系统梳理现有防御策略并指出未解决的关键挑战。
- 适合关注大模型安全与可信AI的研究者阅读。
随着高级推理能力的兴起,大语言模型(LLMs)受到越来越多关注。尽管推理能力提升了下游任务表现,但也引入了新的安全风险:攻击者可利用这些能力实施后门攻击。现有关于后门攻击和推理安全的综述虽全面,但缺乏对针对LLM推理能力的攻击与防御的深入分析。本文首次系统性地综述了基于推理的后门攻击,分析其底层机制、方法框架及未解难题。我们提出一种新分类法,将推理型后门攻击分为关联型、被动型和主动型,提供统一视角。同时,总结防御策略,讨论当前挑战与未来研究方向。本工作为构建安全可信的LLM生态提供新视角。
原文摘要 · Abstract (English)
With the rise of advanced reasoning capabilities, large language models (LLMs) are receiving increasing attention. However, although reasoning improves LLMs' performance on downstream tasks, it also introduces new security risks, as adversaries can exploit these capabilities to conduct backdoor attacks. Existing surveys on backdoor attacks and reasoning security offer comprehensive overviews but lack in-depth analysis of backdoor attacks and defenses targeting LLMs' reasoning abilities. In this paper, we take the first step toward providing a comprehensive review of reasoning-based backdoor attacks in LLMs by analyzing their underlying mechanisms, methodological frameworks, and unresolved challenges. Specifically, we introduce a new taxonomy that offers a unified perspective for summarizing existing approaches, categorizing reasoning-based backdoor attacks into associative, passive, and active. We also present defense strategies against such attacks and discuss current challenges alongside potential directions for future research. This work offers a novel perspective, paving the way for further exploration of secure and trustworthy LLM communities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。