arXiv:2509.04372stat.MLcs.GL2025-09

揭示强化学习与测试时扩展的内在联系,统一多种对齐技术。

Connections between reinforcement learning with feedback,test-time scaling, and diffusion guidance: An anthology

  • 通过反馈机制统一人类与内部反馈的强化学习
  • 发现测试时扩展与扩散引导存在本质关联
  • 提出无需显式强化学习的重采样对齐方法

本文探讨了多种后训练技术之间的深层联系。我们阐明了基于人类反馈的强化学习、基于内部反馈的强化学习与测试时扩展(特别是软最佳N选一采样)之间的紧密关联与等价性,同时揭示了扩散引导与测试时扩展之间的内在联系。此外,我们提出一种用于对齐和奖励导向扩散模型的重采样方法,避免了对显式强化学习技术的需求。

原文摘要 · Abstract (English)

In this note, we reflect on several fundamental connections among widely used post-training techniques. We clarify some intimate connections and equivalences between reinforcement learning with human feedback, reinforcement learning with internal feedback, and test-time scaling (particularly soft best-of-$N$ sampling), while also illuminating intrinsic links between diffusion guidance and test-time scaling. Additionally, we introduce a resampling approach for alignment and reward-directed diffusion models, sidestepping the need for explicit reinforcement learning techniques.

强化学习扩散模型对齐技术测试时扩展

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。