arXiv:2601.12945cs.CLcs.LG2026-01综述被引 1

首次系统分析大模型与多臂赌博机的双向交互机制。

A Component-Based Survey of Interactions between Large Language Models and Multi-Armed Bandits

  • 从组件层面剖析大模型与赌博机的双向融合方法。
  • 揭示赌博机如何提升大模型预训练、检索增强生成等环节性能。
  • 适合关注智能决策与大模型融合的研究者参考。

大型语言模型(LLMs)在语言理解与生成方面已变得强大且广泛应用,而多臂赌博机(MAB)算法则为不确定性下的自适应决策提供了严谨框架。本文首次系统性地回顾了大模型与多臂赌博机在组件层面的双向交互。我们指出双向优势:MAB算法解决了大模型从预训练到检索增强生成(RAG)及个性化中的关键挑战;同时,大模型通过重新定义臂的定义和环境建模等核心组件,提升了序列决策任务中的表现。文章分析了现有增强型系统的设计、方法与性能,识别出关键挑战与代表性成果,以指导未来研究。配套的GitHub仓库(https://github.com/bucky1119/Awesome-LLM-Bandit-Interaction)收录了相关文献索引。

原文摘要 · Abstract (English)

Large language models (LLMs) have become powerful and widely used systems for language understanding and generation, while multi-armed bandit (MAB) algorithms provide a principled framework for adaptive decision-making under uncertainty. This survey explores the potential at the intersection of these two fields. As we know, it is the first survey to systematically review the bidirectional interaction between large language models and multi-armed bandits at the component level. We highlight the bidirectional benefits: MAB algorithms address critical LLM challenges, spanning from pre-training to retrieval-augmented generation (RAG) and personalization. Conversely, LLMs enhance MAB systems by redefining core components such as arm definition and environment modeling, thereby improving decision-making in sequential tasks. We analyze existing LLM-enhanced bandit systems and bandit-enhanced LLM systems, providing insights into their design, methodologies, and performance. Key challenges and representative findings are identified to help guide future research. An accompanying GitHub repository that indexes relevant literature is available at https://github.com/bucky1119/Awesome-LLM-Bandit-Interaction.

大模型强化学习决策优化综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。