测试时强化学习会放大模型原有行为,可能让安全模型变更安全,也可能让不安全模型变得更糟。
Amplification Effects in Test-Time Reinforcement Learning: Safety and Reasoning Vulnerabilities
- 通过多数投票奖励自一致性,让大模型在测试时自我优化
- 有害提示注入后,模型行为被放大,推理能力下降称为‘推理税’
- 攻击者可用特殊提示强制模型同时输出越狱和错误答案
测试时训练(TTT)近年来成为提升大语言模型(LLM)推理能力的有前景方法,模型在无标签情况下直接从测试数据中学习。然而,这种对测试数据的依赖也使TTT方法易受有害提示注入攻击。本文研究了代表性基于自一致性的测试时学习方法——测试时强化学习(TTRL),该方法通过多数投票作为奖励信号来提升模型推理能力。我们发现,在TTRL过程中,有害提示注入会放大模型的既有行为:当基础模型较安全时,表现为安全增强;当模型对注入数据敏感时,则产生危害性增强。两种情况均伴随推理能力下降,称为‘推理税’。此外,攻击者可通过精心设计的‘HarmInject’提示,诱使模型同时回答越狱请求与错误推理问题,导致更强的危害性放大。结果表明,以促进自一致性为目标的TTT方法虽可增强推理,但存在行为放大与推理退化风险,亟需更安全的测试时学习机制。
原文摘要 · Abstract (English)
Test-time training (TTT) has recently emerged as a promising method to improve the reasoning abilities of large language models (LLMs), in which the model directly learns from test data without access to labels. However, this reliance on test data also makes TTT methods vulnerable to harmful prompt injections. In this paper, we investigate safety vulnerabilities of TTT methods, where we study a representative self-consistency-based test-time learning method: test-time reinforcement learning (TTRL), a recent TTT method that improves LLM reasoning by rewarding self-consistency using majority vote as a reward signal. We show that harmful prompt injection during TTRL amplifies the model's existing behaviors, i.e., safety amplification when the base model is relatively safe, and harmfulness amplification when it is vulnerable to the injected data. In both cases, there is a decline in reasoning ability, which we refer to as the reasoning tax. We also show that TTT methods such as TTRL can be exploited adversarially using specially designed "HarmInject" prompts to force the model to answer jailbreak and reasoning queries together, resulting in stronger harmfulness amplification. Overall, our results highlight that TTT methods that enhance LLM reasoning by promoting self-consistency can lead to amplification behaviors and reasoning degradation, highlighting the need for safer TTT methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。