arXiv:2508.04216cs.LG2025-08AAAI被引 4

用因果矫正方法解决推理模型奖励作弊问题,提升数学题解题准确率。

Causal Reward Adjustment: Mitigating Reward Hacking in External Reasoning via Backdoor Correction

  • 通过稀疏自编码器提取奖励模型内部特征,识别干扰性语义因素
  • 在数学推理数据集上使最终准确率提升12.3%,有效减少错误高分路径
  • 无需修改策略模型或重训奖励模型,适合作为通用后处理方案

外部推理系统将语言模型与过程奖励模型(PRM)结合,用于选择复杂任务(如数学问题求解)中的高质量推理路径。然而,这类系统易受奖励作弊影响:逻辑错误但高分的推理路径被错误赋予高分,导致错误答案。从因果推断视角看,该现象主要由混淆语义特征引起。为此,本文提出因果奖励调整(CRA),通过估计推理路径的真实奖励来缓解奖励作弊。CRA 在 PRM 内部激活值上训练稀疏自编码器,恢复可解释特征,并利用后门调整纠正混淆因素。在数学求解数据集上的实验表明,CRA 能有效缓解奖励作弊,提升最终准确率,且无需修改策略模型或重新训练 PRM。

原文摘要 · Abstract (English)

External reasoning systems combine language models with process reward models (PRMs) to select high-quality reasoning paths for complex tasks such as mathematical problem solving. However, these systems are prone to reward hacking, where high-scoring but logically incorrect paths are assigned high scores by the PRMs, leading to incorrect answers. From a causal inference perspective, we attribute this phenomenon primarily to the presence of confounding semantic features. To address it, we propose Causal Reward Adjustment (CRA), a method that mitigates reward hacking by estimating the true reward of a reasoning path. CRA trains sparse autoencoders on the PRM's internal activations to recover interpretable features, then corrects confounding by using backdoor adjustment. Experiments on math solving datasets demonstrate that CRA mitigates reward hacking and improves final accuracy, without modifying the policy model or retraining PRM.

因果推理奖励模型数学推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。