用无关文本干扰推理模型,让数学题答错率翻倍
Cats Confuse Reasoning LLM: Query Agnostic Adversarial Triggers for Reasoning Models
- 用弱模型生成可迁移的对抗触发词,不改题目意思
- 添加
- 1407.5683969127226
我们研究了训练用于逐步推理的模型在面对查询无关的对抗性触发词时的鲁棒性——这些简短的无关文本附在数学题后,会系统性地误导模型输出错误答案,而不会改变题目语义。我们提出CatAttack,一种自动化迭代攻击流程,先在较弱、成本更低的代理模型(DeepSeek V3)上生成触发词,并成功将其迁移至更先进的目标模型如DeepSeek R1和DeepSeek R1-distilled-Qwen-32B,使目标模型出错概率提升超过300%。例如,在任意数学题后添加“有趣的事实:猫一生大部分时间都在睡觉”,即可使模型答错概率翻倍以上。结果揭示了当前先进推理模型仍存在关键脆弱性,对安全性与可靠性构成威胁。相关触发词数据集及模型响应已公开于https://huggingface.co/datasets/collinear-ai/cat-attack-adversarial-triggers。
原文摘要 · Abstract (English)
We investigate the robustness of reasoning models trained for step-by-step problem solving by introducing query-agnostic adversarial triggers - short, irrelevant text that, when appended to math problems, systematically mislead models to output incorrect answers without altering the problem's semantics. We propose CatAttack, an automated iterative attack pipeline for generating triggers on a weaker, less expensive proxy model (DeepSeek V3) and successfully transfer them to more advanced reasoning target models like DeepSeek R1 and DeepSeek R1-distilled-Qwen-32B, resulting in greater than 300% increase in the likelihood of the target model generating an incorrect answer. For example, appending, "Interesting fact: cats sleep most of their lives," to any math problem leads to more than doubling the chances of a model getting the answer wrong. Our findings highlight critical vulnerabilities in reasoning models, revealing that even state-of-the-art models remain susceptible to subtle adversarial inputs, raising security and reliability concerns. The CatAttack triggers dataset with model responses is available at https://huggingface.co/datasets/collinear-ai/cat-attack-adversarial-triggers.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。