arXiv:2509.03456stat.MLcs.LG2025-09

大动作空间下,优化比估计更重要。

Off-Policy Learning in Large Action Spaces: Optimization Matters More Than Estimation

  • 用简化加权似然目标替代复杂估计器,改善优化难题。
  • 在大规模动作空间中,新方法政策性能更优且训练更稳定。
  • 适合研究离线强化学习与大规模决策系统的学者参考。

离线上下文老虎机中的离线策略评估(OPE)和离线策略学习(OPL)是决策的核心。近年来的OPL研究主要聚焦于改进OPE估计器的统计性质,假设更优估计器自然带来更好策略。尽管理论上合理,这种以估计器为中心的方法忽略了实际中关键的挑战:复杂的优化景观。本文提供理论分析与实证证据,表明当前OPL方法在动作空间增大时面临严重优化问题。我们证明,基于估计器感知的策略参数化虽可缓解但无法彻底解决该问题。进一步探索更简单的加权对数似然目标,发现其具备显著更好的优化特性,仍能获得竞争力甚至更优的策略。研究强调,在大规模动作空间中设计OPL算法时,必须显式考虑优化因素。

原文摘要 · Abstract (English)

Off-policy evaluation (OPE) and off-policy learning (OPL) are foundational for decision-making in offline contextual bandits. Recent advances in OPL primarily optimize OPE estimators with improved statistical properties, assuming that better estimators inherently yield superior policies. Although theoretically justified, this estimator-centric approach neglects a critical practical obstacle: challenging optimization landscapes. In this paper, we provide theoretical insights and empirical evidence showing that current OPL methods encounter severe optimization issues, particularly as the action space grows. We show that estimator-aware policy parametrization can mitigate, but not fully resolve, optimization challenges. Building on this, we explore simpler weighted log-likelihood objectives and demonstrate that they enjoy substantially better optimization properties and still recover competitive, often superior, learned policies. Our findings emphasize the necessity of explicitly addressing optimization considerations in the development of OPL algorithms for large action spaces.

离线学习优化难题动作空间策略学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。