arXiv:2511.10686cs.CL2025-11被引 1

研究提示词微调对大模型攻击成功率的影响,发现微小改动可显著改变漏洞暴露程度。

A methodological analysis of prompt perturbations and their effect on attack success rates

  • 对比SFT、DPO、RLHF三种对齐方法,分析提示词扰动对攻击成功率的影响。
  • 统计测试显示,微小提示词变化可导致攻击成功率大幅波动。
  • 提醒仅靠现有攻击基准不足以发现全部漏洞,适合安全评估与模型审计者阅读。

本研究旨在探究不同大型语言模型(LLM)对齐方法如何影响模型对提示词攻击的响应。我们选取基于最常见对齐方法的开源模型,包括监督微调(SFT)、直接偏好优化(DPO)和基于人类反馈的强化学习(RLHF)。通过统计方法进行系统性分析,验证在设计用于诱导不当内容的提示词上施加变化时,攻击成功率(ASR)的敏感性。结果表明,即使微小的提示词修改也会显著改变攻击成功率,使模型更易或更难受到特定攻击。关键的是,研究发现仅运行现有的‘攻击基准’可能不足以揭示模型及对齐方法的所有潜在漏洞。本文通过系统化且基于统计的分析,为模型攻击评估提供了新视角。

原文摘要 · Abstract (English)

This work aims to investigate how different Large Language Models (LLMs) alignment methods affect the models' responses to prompt attacks. We selected open source models based on the most common alignment methods, namely, Supervised Fine-Tuning (SFT), Direct Preference Optimization (DPO), and Reinforcement Learning with Human Feedback (RLHF). We conducted a systematic analysis using statistical methods to verify how sensitive the Attack Success Rate (ASR) is when we apply variations to prompts designed to elicit inappropriate content from LLMs. Our results show that even small prompt modifications can significantly change the Attack Success Rate (ASR) according to the statistical tests we run, making the models more or less susceptible to types of attack. Critically, our results demonstrate that running existing 'attack benchmarks' alone may not be sufficient to elicit all possible vulnerabilities of both models and alignment methods. This paper thus contributes to ongoing efforts on model attack evaluation by means of systematic and statistically-based analyses of the different alignment methods and how sensitive their ASR is to prompt variation.

大模型安全提示词攻击对齐方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。