通过微小权重翻转,悄悄改变大模型的决策立场。
Decision-Level Hijacking: Injecting Cognitive Bias into Large Language Models via Bit-Flip Attacks

- 用可微分情感评估器将主观偏好转为优化信号,实现精准干预。
- 仅翻转少量权重位,就能在多个模型上稳定引发立场偏移。
- 适合关注安全与伦理的AI研究者、系统部署方和政策制定者。
大型语言模型(LLMs)被广泛应用于企业战略等高风险决策场景,用户日益依赖其输出。然而,开源模型共享生态与关键决策应用的深度融合带来了新威胁:若攻击者能操纵模型的认知立场,便可间接影响下游决策者的判断与行为。本文将此类威胁定义为“决策级劫持”。现有攻击无法在不触发禁止内容或降低功能性的前提下实现目标认知操控。为此,本文揭示比特翻转攻击(Bit-Flip Attacks, BFAs)可作为诱导决策级劫持的攻击向量,无需实时交互或控制训练过程,仅需在部署后翻转极少数权重位,即可实现隐蔽、低成本、持久的认知操控。因此,我们提出CogBias框架,通过可微分情感评估器将主观偏好转化为优化信号,采用多目标损失联合约束多维度,构建BitScout定位关键权重位,在极稀疏翻转预算下实现精准认知干预。在Llama-3.2-3B、Mistral-7B和Qwen2.5-14B上的实验,以及商业推荐与争议性事实话题场景验证表明,仅翻转少量比特即可稳定引发目标主题的显著立场偏移,而对非目标任务与整体输出分布影响有限。本工作证明,低层权重数据的微小扰动足以破坏高层价值对齐。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have been widely applied in high-stakes decision-making scenarios such as corporate strategy, and users are increasingly relying on their outputs. However, the deep integration of open-source model sharing ecosystems with LLM-powered critical decision-making applications also introduces critical risks: if an attacker can manipulate the model's cognitive stance, they can indirectly influence the judgments and actions of downstream decision-makers. This paper defines such threats as decision-level hijacking. Existing attacks fail to achieve targeted cognitive manipulation without triggering prohibited content or degrading model functionality. To fill this gap, this paper reveals that Bit-Flip Attacks (BFAs) can serve as an attack vector for inducing decision-level hijacking, requiring no real-time interaction or control over the training process, and only a minimal number of weight bits need to be flipped after deployment to achieve stealthy, low-cost, and persistent cognitive manipulation. Therefore, we propose CogBias, a cognitive bias injection framework for LLMs. CogBias converts subjective preferences into optimization signals via a differentiable sentiment evaluator, uses a multi-objective loss to jointly constrain multiple dimensions, and constructs BitScout to locate critical bits, achieving targeted cognitive intervention under an ultra-sparse flip budget. Experiments on Llama-3.2-3B, Mistral-7B, and Qwen2.5-14B, as well as on the commercial recommendation and controversial factual topic scenarios, demonstrate that flipping only a small number of bits stably induces significant stance shifts on target topics, while the impact on non-target tasks and overall output distribution is limited. This work demonstrates that minute perturbations to low-level weight data suffice to undermine the high-level value alignment of LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。