arXiv:2602.17837cs.CRcs.CL2026-02被引 4

精准操控大模型输出,用少于50次位翻转实现靶向攻击

TFL: Targeted Bit-Flip Attack on Large Language Model

  • 设计关键词聚焦的损失函数,引导生成指定目标词
  • 仅需少于50次位翻转即可成功靶向篡改输出
  • 对无关输入影响极小,适合隐蔽式攻击场景

大型语言模型(LLMs)在安全与关键应用中日益普及,引发对其参数故障注入攻击鲁棒性的担忧。近期研究显示,利用内存漏洞进行位翻转攻击(BFAs)可严重扰乱模型行为。然而,现有方法多导致非靶向失败或普遍性能下降,难以精确控制特定输出。本文提出TFL框架,通过新型关键词聚焦攻击损失函数,促进生成目标词,结合辅助效用评分平衡攻击效果与良性数据的副作用。在Qwen、DeepSeek、Llama等多模型及DROP、GSM8K、TriviaQA等基准上评估,TFL仅用少于50次位翻转即可实现有效靶向输出操纵,且对无关查询影响显著低于以往方法,验证了其高效性与隐蔽性,标志着一类新型精准、低扰动的攻击范式。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly deployed in safety and security critical applications, raising concerns about their robustness to model parameter fault injection attacks. Recent studies have shown that bit-flip attacks (BFAs), which exploit computer main memory (i.e., DRAM) vulnerabilities to flip a small number of bits in model weights, can severely disrupt LLM behavior. However, existing BFA on LLM largely induce un-targeted failure or general performance degradation, offering limited control over manipulating specific or targeted outputs. In this paper, we present TFL, a novel targeted bit-flip attack framework that enables precise manipulation of LLM outputs for selected prompts while maintaining almost no or minor degradation on unrelated inputs. Within our TFL framework, we propose a novel keyword-focused attack loss to promote attacker-specified target tokens in generative outputs, together with an auxiliary utility score that balances attack effectiveness against collateral performance impact on benign data. We evaluate TFL on multiple LLMs (Qwen, DeepSeek, Llama) and benchmarks (DROP, GSM8K, and TriviaQA). The experiments show that TFL achieves successful targeted LLM output manipulations with less than 50 bit flips and significantly reduced effect on unrelated queries compared to prior BFA approaches. This demonstrates the effectiveness of TFL and positions it as a new class of stealthy and targeted LLM model attack.

模型攻击位翻转大模型安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。