arXiv:2605.06785cs.LGcs.AI2026-05

用条件最优传输校准奖励模型,提升成功率预测可靠性。

Distributional Process Reward Models: Calibrated Prediction of Future Rewards via Conditional Optimal Transport

论文配图:Distributional Process Reward Models: Calibrated Prediction of Future Rewards via Conditional Optimal Transport
图 1 · 摘自论文原文
  • 引入条件最优传输学习单调分位数函数,改进奖励模型校准。
  • 在MATH-500和AIME上显著提升校准度,优于未校准模型和分位数回归。
  • 支持任意置信水平的置信区间,适合需要可靠不确定性的推理任务。

推理时缩放方法依赖于过程奖励模型(PRM),但现有方法常存在校准不足、高估成功概率的问题。本文首次提出使用条件最优传输(CondOT)校准PRM,将监督学习中的条件最优传输方法改造为基于PRM隐藏状态估计单调条件分位数函数,从而获得结构合理且可高效提取任意置信水平置信区间的分位数估计。该方法被整合进实例自适应缩放(IAS)框架。在涵盖中等难度问题(MATH-500)和更难分布外问题(AIME)的数学推理基准上测试,对于具有可靠排序信号的PRM,本方法在校准性上显著优于未校准的PRM与分位数回归。下游最佳-选择N的IAS性能也普遍优于未校准的PRM。结果表明,条件最优传输是一种兼具理论保障与实用价值的新型校准范式。

原文摘要 · Abstract (English)

Inference-time scaling methods rely on Process Reward Models (PRMs), which are often poorly calibrated and overestimate success probabilities. We propose, to our knowledge, the first use of conditional optimal transport for calibrating PRMs, modifying conditional OT (CondOT) map learning \cite{bunne2022supervised} to estimate a monotonic conditional quantile function over success probabilities estimated by the PRM, conditioned on PRM hidden states. This yields structurally valid quantile estimates and enables efficient extraction of confidence bounds at arbitrary levels, which we integrate into the instance-adaptive scaling (IAS) framework of \cite{park2025know}. We evaluate on mathematical reasoning benchmarks spanning moderate-difficulty problems (MATH-500) and harder out-of-distribution problems (AIME). For PRMs with reliable ranking signals, our method substantially improves calibration over both uncalibrated PRMs and quantile regression. On downstream Best-of-N IAS performance, our method generally improves over uncalibrated PRMs. These results establish conditional optimal transport as another principled and practical approach to PRM calibration, offering structural guarantees and flexible uncertainty estimation.

奖励模型校准最优传输不确定性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。