微小位翻转可操控生成描述语义,且能用梯度预测关键位。
How a Bit Becomes a Story: Semantic Steering via Differentiable Fault Injection
- 用梯度估计定位影响语义的关键比特,实现可微故障分析。
- 单个比特翻转可改变图像描述的叙事方向,但保持语法正确。
- 为模型鲁棒性测试与可解释性提供新方法,适合安全与可信AI研究者。
硬件中难以察觉的位翻转(如恶意电路或漏洞)已被证明会使Transformer在非生成任务中变得脆弱。本文首次研究大型语言模型(LLM)用于图像描述时,权重中的低层位级扰动(故障注入)如何影响生成描述的语义,同时保持语法结构完整。以往故障分析方法仅关注分类崩溃或准确率下降,忽略了生成系统在语义和语言层面的影响。在图像描述模型中,一个比特翻转可能微妙地改变视觉特征到词语的映射,从而彻底改变人工智能对世界的叙述。我们提出假设:这种语义漂移并非随机,而是可通过模型梯度进行可微估计。为此,我们设计了可微故障分析框架BLADE(Bit-level Fault Analysis via Differentiable Estimation),利用梯度敏感性估计定位语义关键比特,并通过句级语义-流畅性目标优化选择。目标不仅是破坏输出,更是理解语义在比特层面的编码、分布与可变性,揭示即使不可感知的底层变化也能引导生成式视觉-语言模型的高层语义。该研究还为鲁棒性测试、对抗防御和可解释性AI开辟路径,揭示结构化比特故障如何重塑模型的语义输出。
原文摘要 · Abstract (English)
Hard-to-detect hardware bit flips, from either malicious circuitry or bugs, have already been shown to make transformers vulnerable in non-generative tasks. This work, for the first time, investigates how low-level, bitwise perturbations (fault injection) to the weights of a large language model (LLM) used for image captioning can influence the semantic meaning of its generated descriptions while preserving grammatical structure. While prior fault analysis methods have shown that flipping a few bits can crash classifiers or degrade accuracy, these approaches overlook the semantic and linguistic dimensions of generative systems. In image captioning models, a single flipped bit might subtly alter how visual features map to words, shifting the entire narrative an AI tells about the world. We hypothesize that such semantic drifts are not random but differentiably estimable. That is, the model's own gradients can predict which bits, if perturbed, will most strongly influence meaning while leaving syntax and fluency intact. We design a differentiable fault analysis framework, BLADE (Bit-level Fault Analysis via Differentiable Estimation), that uses gradient-based sensitivity estimation to locate semantically critical bits and then refines their selection through a caption-level semantic-fluency objective. Our goal is not merely to corrupt captions, but to understand how meaning itself is encoded, distributed, and alterable at the bit level, revealing that even imperceptible low-level changes can steer the high-level semantics of generative vision-language models. It also opens pathways for robustness testing, adversarial defense, and explainable AI, by exposing how structured bit-level faults can reshape a model's semantic output.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。