arXiv:2409.09819cs.LG2024-09中稿 · the CONSEQUENCES '…

直接优化f散度比原方法更简单有效,提升离线强化学习性能。

A Simpler Alternative to Variational Regularized Counterfactual Risk Minimization

  • 用直接近似替代f-GAN的下界,简化f散度优化流程。
  • 实验表明原f-GAN方法无法复现,新方法性能更优。
  • 适合关注离线强化学习与策略优化的研究者。

变分正则化反事实风险最小化(VRCRM)作为一种离线策略学习(OPL)方法被提出,通过在学习过程中使用日志策略与目标策略间f散度的下界作为正则项,已在多标签分类任务上表现出优于现有方法的性能。本文重新审视VRCRM的原始实验设置,发现其报告结果难以复现。为此,我们提出一种更简单的替代方案:直接最小化f散度的直接近似,而非依赖f-GAN构造的下界。实验显示,基于f-GAN的优化未能达到预期效果,而所提新方法在多个任务上表现更优,具有更强的实用性与可复现性。

原文摘要 · Abstract (English)

Variance regularized counterfactual risk minimization (VRCRM) has been proposed as an alternative off-policy learning (OPL) method. VRCRM method uses a lower-bound on the $f$-divergence between the logging policy and the target policy as regularization during learning and was shown to improve performance over existing OPL alternatives on multi-label classification tasks. In this work, we revisit the original experimental setting of VRCRM and propose to minimize the $f$-divergence directly, instead of optimizing for the lower bound using a $f$-GAN approach. Surprisingly, we were unable to reproduce the results reported in the original setting. In response, we propose a novel simpler alternative to f-divergence optimization by minimizing a direct approximation of f-divergence directly, instead of a $f$-GAN based lower bound. Experiments showed that minimizing the divergence using $f$-GANs did not work as expected, whereas our proposed novel simpler alternative works better empirically.

离线学习策略优化f散度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。