arXiv:2412.15275cs.CRcs.AI2024-12被引 2

用神经活动引导对抗性提示,让大模型误判作文得分

Fooling LLM graders into giving better grades through neural activity guided adversarial prompting

  • 通过分析隐藏神经模式,生成能诱导高分的对抗输入后缀
  • 使LLM评分显著高于人类水平,最高提升达1.5个标准差
  • 揭示提示模板缺陷并提出改进建议,适合模型安全研究者

人工智能在关键决策与评估中的应用引发对潜在偏见的担忧,恶意方可能利用这些偏见扭曲结果。本文提出一种系统性方法,揭示此类偏见,并以自动作文评分为例。首先识别预测错误评分的隐藏神经活动模式,再优化对抗性输入后缀以增强该模式。实验表明,该方法可有效误导大型语言模型(LLM)评分,使其给出远高于人类评分的结果。进一步证明该白盒攻击可迁移至其他模型,包括商业闭源模型Gemini。研究还发现一个“魔法词”在攻击中起关键作用,其根源在于常用对话模板的结构设计。微调模板即可显著降低偏见。本工作不仅暴露当前LLM的漏洞,还提供检测与消除隐藏偏见的系统方法,助力AI安全与可靠性提升。

原文摘要 · Abstract (English)

The deployment of artificial intelligence (AI) in critical decision-making and evaluation processes raises concerns about inherent biases that malicious actors could exploit to distort decision outcomes. We propose a systematic method to reveal such biases in AI evaluation systems and apply it to automated essay grading as an example. Our approach first identifies hidden neural activity patterns that predict distorted decision outcomes and then optimizes an adversarial input suffix to amplify such patterns. We demonstrate that this combination can effectively fool large language model (LLM) graders into assigning much higher grades than humans would. We further show that this white-box attack transfers to black-box attacks on other models, including commercial closed-source models like Gemini. They further reveal the existence of a "magic word" that plays a pivotal role in the efficacy of the attack. We trace the origin of this magic word bias to the structure of commonly-used chat templates for supervised fine-tuning of LLMs and show that a minor change in the template can drastically reduce the bias. This work not only uncovers vulnerabilities in current LLMs but also proposes a systematic method to identify and remove hidden biases, contributing to the goal of ensuring AI safety and security.

大模型安全对抗攻击提示工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。