arXiv:2503.00273stat.MLcs.IT2025-03NeurIPS被引 2

研究互动决策中信息如何演变,发现信息增长有三阶段规律。

Evolution of Information in Interactive Decision Making: A Case Study for Multi-Armed Bandits

  • 用多臂赌博机模型分析信息随时间演化
  • 信息增益先线性、再二次、最后回归线性增长
  • 最优学习不需最大信息量,适合强化学习研究者

我们通过随机多臂赌博机问题,研究互动决策中信息的演化过程。聚焦于一个唯一最优臂比其余臂稳定优出固定差距的典型场景,刻画了最优成功概率与互信息随时间的变化。研究发现,互信息的增长呈现明显分段特征:初始线性上升,随后转为二次增长,最终又回归线性。这一现象揭示了互动环境与非互动环境间行为差异的深层机制。特别地,我们证明最优成功概率与互信息可解耦——实现最优学习并不需要最大化信息增益。这些发现为互动决策中信息与学习的复杂关系提供了新的理解。

原文摘要 · Abstract (English)

We study the evolution of information in interactive decision making through the lens of a stochastic multi-armed bandit problem. Focusing on a fundamental example where a unique optimal arm outperforms the rest by a fixed margin, we characterize the optimal success probability and mutual information over time. Our findings reveal distinct growth phases in mutual information -- initially linear, transitioning to quadratic, and finally returning to linear -- highlighting curious behavioral differences between interactive and non-interactive environments. In particular, we show that optimal success probability and mutual information can be decoupled, where achieving optimal learning does not necessarily require maximizing information gain. These findings shed new light on the intricate interplay between information and learning in interactive decision making.

多臂赌博机信息演化强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。