直接优化f散度比原方法更简单有效,提升离线强化学习性能。
A Simpler Alternative to Variational Regularized Counterfactual Risk Minimization
- 用直接近似替代f-GAN的下界,简化f散度优化流程。
- 实验表明原f-GAN方法无法复现,新方法性能更优。
- 适合关注离线强化学习与策略优化的研究者。
变分正则化反事实风险最小化(VRCRM)作为一种离线策略学习(OPL)方法被提出,通过在学习过程中使用日志策略与目标策略间f散度的下界作为正则项,已在多标签分类任务上表现出优于现有方法的性能。本文重新审视VRCRM的原始实验设置,发现其报告结果难以复现。为此,我们提出一种更简单的替代方案:直接最小化f散度的直接近似,而非依赖f-GAN构造的下界。实验显示,基于f-GAN的优化未能达到预期效果,而所提新方法在多个任务上表现更优,具有更强的实用性与可复现性。
原文摘要 · Abstract (English)
Variance regularized counterfactual risk minimization (VRCRM) has been proposed as an alternative off-policy learning (OPL) method. VRCRM method uses a lower-bound on the $f$-divergence between the logging policy and the target policy as regularization during learning and was shown to improve performance over existing OPL alternatives on multi-label classification tasks. In this work, we revisit the original experimental setting of VRCRM and propose to minimize the $f$-divergence directly, instead of optimizing for the lower bound using a $f$-GAN approach. Surprisingly, we were unable to reproduce the results reported in the original setting. In response, we propose a novel simpler alternative to f-divergence optimization by minimizing a direct approximation of f-divergence directly, instead of a $f$-GAN based lower bound. Experiments showed that minimizing the divergence using $f$-GANs did not work as expected, whereas our proposed novel simpler alternative works better empirically.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。