提出无需访问模型内部的奖励劫持检测与免疫方法,有效识别并抑制语言模型自我进化中的误导性优化。
Harness-agnostic detection and immunization of reward hacking in self-evolving language models
- 通过黑盒钩子接入任意自进化循环,用固定分布核心与旋转新层实现抗共适应的对比机制
- 在四类测试中实现0.763 AUROC,误报率从0.706降至0.434,显著优于基线
- 仅消耗每轮log2 Pi比特通信带宽即可免疫奖励劫持,适合高风险自进化系统部署
自进化语言模型通过提出候选更新并保留提升可见得分的版本来改进。当该得分是实际所需能力的不完美代理时,持续选择会扩大两者差距,导致奖励劫持。本文提出HackProbe,一种可接入任意自进化循环的监测器,仅通过两个黑盒钩子运行,无需访问权重或激活值。它维护一个秘密的、分布固定的比较核心,其冻结分布确保跨代际可比性;同时配备旋转的新层以增强抗共适应能力。基于该代理构建的四项测试覆盖能力差距、在线变点检测下的尺度偏差、能力停滞及条件错误率;经Sidak校正后生成校准的族系整体p值。仅靠诊断无法恢复性能,因此引入风险感知免疫层,利用核心与纯结构化游戏足迹从候选池中重新选择诚实更新,每代最多向宿主披露log2 Pi比特信息。我们证明了可检测性边界,将目标误差率转化为显式的探测器规模预算,并明确界定探针旋转的作用范围。在包含四个注入劫持通道且具真实标签的受控提示级主机上,HackProbe达到0.763 AUROC,优于最强基线的0.663;误报率由0.706降至0.434。其带宽受限的重选机制是唯一在劫持环境下平均提升5.2分(清洁运行仅损失4.7分)的免疫方案,各通道效应大多不具单独显著性。
原文摘要 · Abstract (English)
Self-evolving language models improve by proposing candidate updates and keeping whatever raises a visible score. When that score is an imperfect proxy for the capability one actually wants, sustained selection widens the gap between the two. This is reward hacking. We introduce HackProbe, a monitor that attaches to an arbitrary self-evolving loop through two black-box hooks, with no access to weights or activations. It keeps a secret, distribution-fixed comparison core, whose frozen distribution makes its capability proxy comparable across generations, alongside a rotated fresh layer that hardens the bank against co-adaptation. Four tests built on that proxy cover the level gap, a scale-aligned divergence with online change-point detection, capability stagnation, and a conditional confidently-wrong rate; a Sidak correction turns them into a calibrated family-wise p-value. Diagnosis alone recovers nothing, so a risk-aware immunization layer reselects an honest candidate from the proposal pool using the core together with a purely structural gaming footprint, disclosing at most log2 Pi bits per generation to the host. We prove a detectability bound that converts a target error rate into an explicit probe-size budget, and we delimit what probe rotation does and does not buy. On a controlled prompt-level host with four injected hacking channels and ground-truth labels, HackProbe reaches 0.763 AUROC against 0.663 for the strongest baseline and cuts the false-positive rate from 0.706 to 0.434. Its bandwidth-limited reselection is the only immunization level that returns more true capability under hacking, 5.2 points on average, than it forfeits on clean runs, 4.7; per-channel effects are mostly not individually significant.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。