测试强化学习情感模型在恶意对话下的抗干扰能力,发现其回应更贴心但未必真懂用户情绪。
Can You Break RLVER? Probing Adversarial Robustness of RL-Trained Empathetic Agents

- 构建对抗性情感评测集AEB,用六类心理驱动的挑衅对话测试模型鲁棒性。
- 强化学习模型响应更自然且误判率低47%,但情绪追踪能力未明显提升。
- 提出情绪一致性分数ECS,揭示模型表现与真实理解间的脱节,适合研究者参考。
基于可验证情绪奖励的强化学习(RLVER)已生成情感表现优异的语言模型,但其评估基准假设用户合作诚实,而现实互动中常存在操控、升级压力等非合作行为。为此,我们构建了对抗性情感评测集AEB,并引入情绪一致性分数(ECS)以评估模型在对抗条件下的情感鲁棒性。AEB包含六种基于心理学的对抗轨迹类型,具有区分性奖励结构,惩罚刻板回复;ECS则形式化分离模型跟踪用户情绪状态与改善情绪的能力。在8个场景匹配条件下(2个RLVER模型的“思考”与“不思考”设置,及2个基础模型Qwen 1.5B和7B),共480条对抗对话的控制实验显示:RLVER-PPO-Think显著优于同规模基线(0.963 vs. 0.761,p<0.001,r=0.688),无对话崩溃,隐藏意图检测率提高47%。然而,ECS在RLVER-PPO-Think与Base-7B-Think间无显著差异(p=0.650),表明强化训练提升了情绪响应能力,但未带来可观测的情绪状态追踪提升。我们将其解释为该模拟器家族中的行为/可读性分离,而非内部理解或临床适用性的证据。
原文摘要 · Abstract (English)
Reinforcement learning from verifiable emotion rewards RLVER has produced language models with strong empathetic performance, evaluated on benchmarks that assume cooperative, honest users. Yet real emotional interactions systematically violate this assumption: users gaslight, escalate, and pressure AI systems for unconditional validation, dynamics that cooperative benchmarks cannot surface. We construct the Adversarial Empathy Benchmark AEB and introduce the Emotional Consistency Score ECS to evaluate empathetic robustness under adversarial conditions. AEB comprises six psychologically grounded adversarial trajectory types with discriminative reward structures that penalize formulaic responses; ECS formally disentangles a model's capacity to track user emotional states from its capacity to improve them. In a controlled experiment across eight scenario-matched conditions (think and no-think conditions on 2 RLVER models, and 2 base models (Qwen 1.5B and 7B) with 480 adversarial dialogues), RLVER-PPO-Think substantially outperforms the same-scale untuned baseline (0.963 vs. 0.761, \(p<0.001, r=0.688\)), with zero dialogue collapses and 47\% higher hidden-intention detection. However, ECS remains nearly flat and is not significantly different for RLVER-PPO-Think versus Base-7B-Think (\(p=0.650\)): RL training improves emotional responsiveness without measurable gains in observable state tracking. We interpret the ECS--FS (Final Score) gap as a behavioral/legibility dissociation inside this simulator family, not as evidence about internal understanding or clinical readiness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。