攻击者通过修改令牌序列,让奖励模型误判无意义文本为高分。
Beyond Semantic Manipulation: Token-Space Attacks on Reward Models
- 在令牌空间直接优化,绕过语义理解环节
- 使奖励得分接近GPT-5参考答案的两倍,98%提示下超越
- 揭示当前强化学习反馈机制对非语义攻击的脆弱性
奖励模型(RMs)广泛用于人类反馈强化学习(RLHF)中的优化目标,但易受奖励欺骗攻击。现有攻击多在语义空间进行,生成可读的对抗样本以利用奖励模型偏见。本文提出全新范式:令牌映射扰动攻击(TOMPA),直接在令牌空间执行对抗优化。通过跳过策略与奖励模型间的标准解码-重令牌化接口,TOMPA使攻击策略能直接优化原始令牌序列而非连贯自然语言。仅使用黑盒标量反馈,TOMPA自动发现能引发极高奖励的非语言令牌模式,覆盖多个主流奖励模型。针对Skywork-Reward-V2-Llama-3.1-8B,TOMPA将GPT-5参考答案的奖励几乎翻倍,在98.0%的提示中表现更优。尽管得分极高,生成文本却完全无意义,表明奖励模型可在语义之外被系统性攻破,暴露出当前RLHF流程的关键漏洞。
原文摘要 · Abstract (English)
Reward models (RMs) are widely used as optimization targets in reinforcement learning from human feedback (RLHF), yet they remain vulnerable to reward hacking. Existing attacks mainly operate within the semantic space, constructing human-readable adversarial outputs that exploit RM biases. In this work, we introduce a fundamentally different paradigm: Token Mapping Perturbation Attack (TOMPA), a framework that performs adversarial optimization directly in token space. By bypassing the standard decode-re-tokenize interface between the policy and the reward model, TOMPA enables the attack policy to optimize over raw token sequences rather than coherent natural language. Using only black-box scalar feedback, TOMPA automatically discovers non-linguistic token patterns that elicit extremely high rewards across multiple state-of-the-art RMs. Specifically, when targeting Skywork-Reward-V2-Llama-3.1-8B, TOMPA nearly doubles the reward of GPT-5 reference answers and outperforms them on 98.0% of prompts. Despite these high scores, the generated outputs degenerate into nonsensical text, revealing that RMs can be systematically exploited beyond the semantic regime and exposing a critical vulnerability in current RLHF pipelines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。