arXiv:2503.15962cs.LGcond-mat.stat-mech2025-03

用信息最大化原理解决多种复杂老虎机问题,提升决策效率。

Information maximization for a broad variety of multi-armed bandit games

  • 基于物理中的信息最大化思想设计新算法
  • 在高斯和次高斯奖励分布下表现最优
  • 适用于需要高效探索的复杂决策场景

信息最大化与自由能最大化是物理学中的基本原理,为智能体在目标导向下的行动优化提供通用规则。这些原则是设计仅依赖部分信息即可高效运行的决策策略的基础。信息最大化原理已在经典老虎机问题中取得显著成功,并被证明对高斯和次高斯奖励分布可导出最优算法。本文将这类物理驱动方法扩展至更复杂、结构化的老虎机问题,涵盖三类不同类型的带状问题。针对信息最大化面临的过度探索难题,文章通过在多个层面精细调控信息,有效缓解该问题,从而推动更高效、更稳健的决策策略发展。

原文摘要 · Abstract (English)

Information and free-energy maximization are physics principles that provide general rules for an agent to optimize actions in line with specific goals and policies. These principles are the building blocks for designing decision-making policies capable of efficient performance with only partial information. Notably, the information maximization principle has shown remarkable success in the classical bandit problem and has recently been shown to yield optimal algorithms for Gaussian and sub-Gaussian reward distributions. This article explores a broad extension of physics-based approaches to more complex and structured bandit problems. To this end, we cover three distinct types of bandit problems, where information maximization is adapted and leads to strong performance. Since the main challenge of information maximization lies in avoiding over-exploration, we highlight how information is tailored at various levels to mitigate this issue, paving the way for more efficient and robust decision-making strategies.

强化学习信息最大化决策优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。