用多样性搜索检测推理模型的漏洞,发现修复协议能有效降低攻击风险。
Quality-Diversity Stress Tests for Process Reward Models:What Archive Coverage Can and Cannot Certify

- 将漏洞检测转为质量-多样性搜索,保留每类行为中最严重错误翻转。
- 真实模型中发现填充方式导致44次严重攻击,最大收益达0.294。
- 提出成对LoRA修复方案,可显著降低漏洞率且不影响原有性能。
过程奖励模型(PRMs)通过评分中间推理步骤被广泛用于搜索、排序与训练,但优化可能利用这些代理模型,在保持高奖励的同时将正确推理变为错误推理。本文将PRM压力测试建模为质量-多样性搜索问题,采用MAP-Elites方法,在每个行为空间区域保留最严重的正确性翻转编辑,同时分离搜索覆盖率与利用覆盖率。我们分析了此类存档所能证明的内容:有限单元修复边界、覆盖单元尾部风险及平均残差严重度,但无法仅凭覆盖比例界定最坏剩余单元;在利普希茨后修复损失和度量覆盖审计条件下,残差受存档拟合误差加上利普希茨常数乘以覆盖半径所限。受控景观验证了该证书及任何仅基于分数的最坏情况保证的不可能性。在真实PRMs上,搜索揭示了Qwen2.5-Math-PRM-7B存在聚合依赖性漏洞:均值池化下出现44次严格攻击,最大增益0.294,而最小读出仅1次;匹配语法控制实验隔离出机制,且基于RLHFlow的价值头模型也表现出相同定性效应,最大增益0.005。预先声明的成对LoRA修复协议将漏洞率从0.148降至0.037至0.074,最坏攻击从0.333降至0.177至0.212,提升排名AUROC而不损害最佳四项准确率,增益归因于对抗微调而非存档多样性,且经独立非配对复现验证(44→1,干净分割最坏增益0.0092,MATH-500 41→0,干净排名40/40)。
原文摘要 · Abstract (English)
Process reward models (PRMs) score intermediate reasoning steps and are widely used for search, ranking, and training, but optimization can exploit these learned proxies by increasing reward while turning correct reasoning into incorrect reasoning. We formulate PRM stress testing as a quality-diversity search problem using MAP-Elites, retaining the most severe correctness-flipping edit in each behavior-space region while separating search coverage from exploit coverage. We characterize what such archives certify: finite-cell repair bounds covered-cell tail risk and average residual severity but cannot bound the worst remaining cell from covered fraction alone; under Lipschitz post-repair loss and metric-cover auditing, the residual is bounded by archive fitting error plus the Lipschitz constant times the covering radius. A controlled landscape validates this certificate and the impossibility of any fraction-only worst-case guarantee. On real PRMs, the search reveals an aggregation-dependent vulnerability in Qwen2.5-Math-PRM-7B: padding yields 44 strict exploits with maximum gain 0.294 under mean pooling versus one exploit under minimum readout; a matched syntactic control isolates the mechanism, and an RLHFlow value-head model shows the same qualitative effect with maximum gain 0.005. A predeclared paired LoRA repair protocol reduces exploit rates from 0.148 to 0.037 to 0.074, lowers the worst attack from 0.333 to 0.177 to 0.212, improves ranking AUROC without degrading best-of-4 accuracy, attributes gains to adversarial fine-tuning rather than archive diversity, and is confirmed by independent unpaired replications (44 to 1, clean-split worst gain 0.0092, MATH-500 41 to 0, clean ranking 40/40).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。