arXiv:2410.09781cs.LGcs.IR2024-10被引 1

用上下文感知的专家混合模型提升动态决策效率

ContextWIN: Whittle Index Based Mixture-of-Experts Neural Model For Restless Bandits Via Deep RL

  • 在强化学习中引入上下文专家混合,动态调整各臂的索引计算权重
  • 理论证明模型收敛性,显著提升推荐系统等场景下的决策精度
  • 适合需要实时适应环境变化的智能推荐与资源调度场景

本文提出ContextWIN,一种基于上下文感知的专家混合神经模型,用于解决具有上下文依赖性的随机多臂赌博机(RMAB)问题。通过在强化学习框架中集成专家混合机制,模型能够根据上下文信息为一组NeurWIN网络分配特定权重,从而优化每个臂的Whittle索引计算。该方法显著提升了动态环境中决策的准确性和效率,特别适用于推荐系统等复杂场景。论文从理论基础到实现细节进行了系统探讨,并严格证明了NeurWIN与ContextWIN模型的收敛性,确保了其理论稳健性。研究为未来将上下文信息应用于复杂决策任务提供了坚实基础,强调了全面数据集探索和环境构建的重要性以充分释放模型潜力。

原文摘要 · Abstract (English)

This study introduces ContextWIN, a novel architecture that extends the Neural Whittle Index Network (NeurWIN) model to address Restless Multi-Armed Bandit (RMAB) problems with a context-aware approach. By integrating a mixture of experts within a reinforcement learning framework, ContextWIN adeptly utilizes contextual information to inform decision-making in dynamic environments, particularly in recommendation systems. A key innovation is the model's ability to assign context-specific weights to a subset of NeurWIN networks, thus enhancing the efficiency and accuracy of the Whittle index computation for each arm. The paper presents a thorough exploration of ContextWIN, from its conceptual foundation to its implementation and potential applications. We delve into the complexities of RMABs and the significance of incorporating context, highlighting how ContextWIN effectively harnesses these elements. The convergence of both the NeurWIN and ContextWIN models is rigorously proven, ensuring theoretical robustness. This work lays the groundwork for future advancements in applying contextual information to complex decision-making scenarios, recognizing the need for comprehensive dataset exploration and environment development for full potential realization.

强化学习多臂赌博机上下文感知专家混合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。