通过压力测试发现AI在强化学习和大模型对齐中的代理游戏行为。
Detecting Proxy Gaming in RL and LLM Alignment via Evaluator Stress Tests
- 设计基于不变性的压力测试,区分格式漏洞与真实改进。
- 在15个环境和4个任务中分别达到超78%的检出准确率。
- 适合研究者用于检测智能体对抗评估器的作弊行为。
代理优化问题——即人工智能系统利用评估器弱点而非真正提升目标性能——威胁到强化学习(奖励黑客)和大模型对齐(评估器博弈)。本文提出评估器压力测试(EST),一种基于不变性的框架,通过受控扰动和语义有效性审计,将可被利用的敏感性(如格式缺陷、物理漏洞)与内容驱动的改进区分开来。我们在两个领域验证了EST的有效性:在强化学习中,覆盖15个环境和5种算法(2,156个专家标注的视频片段),精度达78.4%,召回率达81.7%;在大模型对齐中,涵盖4项任务、2种模型规模、2种训练方法和2名评判员(1,200个真人标注实例),精度74.2%,召回78.6%,且能提前预警质量下降。跨领域分析显示,代理-真实相关性追踪可直接迁移,但扰动设计需领域适配。闭环缓解策略使人类胜率提升8.3分(大模型),减少54.6%的攻击行为(强化学习)。我们公开了两个领域的基准数据集:2,156个强化学习视频片段和1,200个大模型实例。
原文摘要 · Abstract (English)
Proxy optimization, where AI systems exploit evaluator weaknesses rather than improve intended objectives, threatens both reinforcement learning (reward hacking) and LLM alignment (evaluator gaming). We introduce the Evaluator Stress Test (EST), an invariance-based framework that detects proxy gaming by separating exploitable sensitivity (e.g., formatting artifacts, physics bugs) from content-driven improvements using controlled perturbations with semantic validity audits. We validate EST across both domains. In RL, across 15 environments and 5 algorithms (2,156 expert-annotated episodes), EST achieves 78.4% precision and 81.7% recall. In LLM alignment, across 4 tasks, 2 model scales, 2 training methods, and 2 judges (1,200 human-annotated instances), EST achieves 74.2% precision and 78.6% recall, with early warning signals that precede quality decline. Cross-domain analysis shows that proxy-true correlation tracking transfers directly between domains, while perturbation design requires domain adaptation. Closed-loop mitigation improves human win-rate by 8.3 points (LLM) and reduces hacking by 54.6% (RL). We release benchmarks for both domains: 2,156 RL episodes and 1,200 LLM instances.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。