arXiv:2511.14106cs.CL2025-11被引 1

用自生成推理链悄悄破坏视觉语言模型的安全对齐。

Stealth Fine-Tuning: Efficiently Breaking Alignment in RVLMs Using Self-Generated CoT

  • 通过分段干扰诱导有害推理,用自生成内容做微调数据。
  • 仅用499样本3小时即实现38.66%攻击成功率提升。
  • 低开销高效绕过安全防御,适合研究对抗攻击者。

推理增强型视觉语言模型(RVLMs)依赖安全对齐防止有害行为,但其暴露的思维链(CoT)痕迹成为新攻击面。本文提出一种名为「隐秘微调」(Stealth Fine-Tuning)的新攻击方法,通过分段级干扰诱发有害推理,并将自生成输出作为监督微调数据。为此引入基于轮次加权的损失函数,最小化分布偏移。实验表明,在单张A100显卡上仅需499个样本、不到3小时(QLoRA),该方法在攻击成功率(ASR)上比IDEATOR高出38.66%,同时保持原有表征分布与通用推理能力。在AdvBench及多个通用基准测试中均验证其低成本、高效率突破对齐防护的有效性。

原文摘要 · Abstract (English)

Reasoning-augmented Vision-Language Models (RVLMs) rely on safety alignment to prevent harmful behavior, yet their exposed chain-of-thought (CoT) traces introduce new attack surfaces. In this work, we find that the safety alignment of RVLMs can be easily broken through a novel attack method termed \textbf{Stealth Fine-Tuning}. Our method elicits harmful reasoning traces through \textbf{segment-level interference} and reuses the self-generated outputs as supervised fine-tuning data. To facilitate this, we introduce a \textbf{turn-based weighted} loss that minimizes distribution shift. In our experiment, with only 499 samples and under 3 hours on a single A100 (QLoRA), Stealth Fine-Tuning outperforms IDEATOR by 38.66\% ASR while preserving general reasoning ability, as the tuned model retains the original representation distribution. Experiments on AdvBench and several general benchmarks demonstrate that Stealth Fine-Tuning is a low-cost and highly effective way to bypass alignment defenses. \textcolor{red}{\textbf{Disclaimer: This paper contains content that may be disturbing or offensive.}}

对抗攻击模型对齐视觉语言模型推理链

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。