arXiv:2605.02395cs.AI2026-05

通过可控错误注入生成可验证的推理过程标注数据,提升推理模型训练效果。

Verifiable Counterfactual Supervision for Process Reward Models

论文配图:Verifiable Counterfactual Supervision for Process Reward Models
图 1 · 摘自论文原文
  • 在符号推理链中选定步骤注入模板感知错误并重新计算后续步骤
  • 确保错误步骤无法从原始前缀推导,实现首个无效转移的精确标注
  • 生成的合成数据显著提升逻辑推理任务的重排性能,适用于数学推理评估

过程奖励模型(PRMs)需要不仅判断推理路径是否正确,还需定位推理过程中首次脱离前缀支持的位置。本文提出可验证反事实过程监督:基于已验证的符号推理链,在选定中间步骤注入模板感知错误,重新计算后续所有步骤,确保注入步骤无法由原前缀推导得出,且下游延续在污染状态下仍保持连贯性。由此生成的轨迹包含前缀有效性的首错标注,并转化为自然语言过程用于PRM训练与评估。实验表明,合成数据在逻辑推理基准上使Best-of-8重排性能提升,初步展现向数学过程评估的迁移能力。

原文摘要 · Abstract (English)

Process reward models (PRMs) require supervision that identifies not only whether a reasoning trajectory is correct, but also where the reasoning process first becomes unsupported by its prefix. We frame this requirement as verifiable counterfactual process supervision with paired correct and erroneous trajectories in which the first invalid transition is known, the error mechanism is controlled, and the downstream continuation remains coherent under the corrupted state. Starting from a verified symbolic reasoning chain, our method injects a template-aware error at a selected intermediate step, recomputes all subsequent steps under the corrupted state, and verifies that the injected step is not derivable from its original prefix. The resulting trajectories provide prefix-valid first-error annotations and are translated into aligned natural-language processes for PRM training and evaluation. Experiments show that the synthesized data improve Best-of-8 reranking on logical reasoning benchmarks and show preliminary transfer to mathematical process evaluation.

推理建模奖励模型可验证性合成数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。