研究图结构与大规模动作空间中的强化学习问题,提升实际应用可行性。
Bandits on graphs and structures

- 将动作结构建模为图,利用谱平滑性与侧信息优化决策
- 在指数级或无限动作空间中实现高效探索与最优率收敛
- 适合对复杂决策场景建模的研究者与工业应用开发者
本论文旨在研究某些序列决策问题的结构性质,以推动其向实际应用靠拢。第一部分聚焦可表示为动作图的结构,研究奖励平滑性(谱带模型)、侧观察以及影响力最大化等设置。第二部分探讨动作空间规模可达基动作数的指数级甚至无限大的情形,涵盖核函数带模型、多子模带模型、函数优化带模型(含未知光滑性)及无穷臂带模型。本文系统总结了作者在图结构与结构化带模型方面的研究成果。
原文摘要 · Abstract (English)
The goal of this thesis is to investigate the structural properties of certain sequential problems in order to bring the solutions closer to a practical use. In the first part, we put a special emphasis on structures that can be represented as graphs on actions. In the second part, we study the large action spaces that can be of exponential size in the number of base actions or even infinite. For graph bandits, we consider the settings of smoothness of rewards (spectral bandits), side observations, and influence maximization. For large structured domains, we cover kernel bandits, polymatroid bandits, bandits for function optimization (including unknown smoothness), and infinitely many-arms bandits. The thesis aspires to be a survey of the author's contributions on graph and structured bandits.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。