用协方差几何分析评分模型集成的漏洞,给出有限搜索下的可靠保证。
Evaluator Ensembles Under Reward Hacking: Covariance Geometry and Finite-Search Guarantees

- 通过评分器协方差结构揭示共模误差与分歧的关系。
- 证明仅凭内部评分无法识别共模误差,且可量化最优选择的过估计。
- 适合研究评分模型鲁棒性与强化学习对齐的学者参考。
语言模型评分器和奖励模型实现可扩展监督,但有限优化可能利用评分器错误而非提升响应质量。本文通过评分器集成的协方差几何刻画此失败现象。对于校准的评分器,集成均值保留沿全一向量方向的共模误差,而评分器间分歧仅捕捉正交误差。因此,分歧可能很高但聚合仍稳健,或很低但共享的响应依赖误差持续存在。我们证明共模误差无法仅从内部评分中识别。在联合次高斯模型下,我们界定了最佳-K选择的过估计和目标质量遗憾,将保证扩展至条件校准下的可预测自适应搜索。所得搜索项随√log K 缩放,对高斯投影误差渐近紧致。进一步表明,噪声质量代理引入非真实秩一协方差而不改变分歧,并提出有界双锚伯恩斯坦证书以控制有限搜索误差与遗憾。对120组 (J, ρ, K) 配置的固定种子高斯压力测试及真实模型审计验证了理论,同时揭示分歧诊断在增加搜索压力下的局限性。
原文摘要 · Abstract (English)
Language-model judges and reward models enable scalable supervision, but finite optimization can exploit evaluator errors rather than improve response quality. We characterize this failure through the covariance geometry of evaluator ensembles. For calibrated judges, the ensemble mean retains common-mode error along the all-ones direction, whereas cross-judge disagreement captures only orthogonal error. Consequently, disagreement can be high despite robust aggregation, or low while shared response-dependent errors persist. We prove that common-mode error is not identifiable from internal judge scores alone. Under a joint sub-Gaussian model, we bound best-of-K selection overstatement and target-quality regret, extending the guarantees to predictably adaptive search under conditional calibration. The resulting search terms scale as the square root of log K and are asymptotically tight for Gaussian projected errors. We further show that noisy quality proxies introduce artificial rank-one covariance without changing disagreement, and propose a bounded two-anchor Bernstein certificate for finite-search error and regret. Fixed-seed Gaussian stress tests over 120 (J, rho, K) configurations and real-model audits validate the theory while revealing the limits of disagreement-based diagnostics under increasing search pressure.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。