arXiv:2510.22686cs.LGcs.AI2025-10TPAMI被引 4

用流匹配生成价值分布,提升强化学习中价值估计的准确性与表达能力。

FlowCritic: Bridging Value Estimation with Flow Matching in Reinforcement Learning

  • 基于流匹配建模价值分布,生成样本进行价值估计
  • 相比传统点估计,能更好捕捉复杂价值分布特性
  • 适合追求高精度价值评估的强化学习研究者

可靠的值函数估计是强化学习的基石,用于评估长期回报并指导策略改进,显著影响收敛速度和最终性能。现有方法通过多评论家集成和分布强化学习提升估计可靠性,但前者仅组合多个点估计而未捕捉分布信息,后者依赖离散化或分位数回归,限制了复杂价值分布的表达能力。受生成建模中流匹配成功的启发,我们提出一种生成式价值估计范式——FlowCritic。不同于传统的确定性值预测回归方法,FlowCritic利用流匹配建模价值分布,并生成样本以实现价值估计。

原文摘要 · Abstract (English)

Reliable value estimation serves as the cornerstone of reinforcement learning (RL) by evaluating long-term returns and guiding policy improvement, significantly influencing the convergence speed and final performance. Existing works improve the reliability of value function estimation via multi-critic ensembles and distributional RL, yet the former merely combines multi point estimation without capturing distributional information, whereas the latter relies on discretization or quantile regression, limiting the expressiveness of complex value distributions. Inspired by flow matching's success in generative modeling, we propose a generative paradigm for value estimation, named FlowCritic. Departing from conventional regression for deterministic value prediction, FlowCritic leverages flow matching to model value distributions and generate samples for value estimation.

强化学习价值估计流匹配生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。