arXiv:2505.13355cs.AI2025-05AAAI综述被引 9

用强化学习优化大模型,让大模型反哺决策算法。

Survey: Multi-Armed Bandits Meet Large Language Models

  • 用多臂赌博机平衡大模型训练中的探索与利用。
  • 大模型通过自然语言推理提升决策算法的适应性。
  • 适合关注智能决策与大模型融合的研究者。

多臂赌博机算法与大语言模型(LLMs)是人工智能中两个强大的工具,分别解决决策优化与自然语言处理问题。本文综述两者结合的潜力:赌博机可优化大模型的微调、提示工程与自适应生成,通过平衡探索与利用提升大规模学习效率;而大模型则通过上下文理解、动态适应和自然语言推理,增强赌博机的策略选择能力。文章系统梳理现有研究,指出关键挑战与未来机遇,旨在弥合两领域鸿沟,推动人工智能跨学科创新。

原文摘要 · Abstract (English)

Bandit algorithms and Large Language Models (LLMs) have emerged as powerful tools in artificial intelligence, each addressing distinct yet complementary challenges in decision-making and natural language processing. This survey explores the synergistic potential between these two fields, highlighting how bandit algorithms can enhance the performance of LLMs and how LLMs, in turn, can provide novel insights for improving bandit-based decision-making. We first examine the role of bandit algorithms in optimizing LLM fine-tuning, prompt engineering, and adaptive response generation, focusing on their ability to balance exploration and exploitation in large-scale learning tasks. Subsequently, we explore how LLMs can augment bandit algorithms through advanced contextual understanding, dynamic adaptation, and improved policy selection using natural language reasoning. By providing a comprehensive review of existing research and identifying key challenges and opportunities, this survey aims to bridge the gap between bandit algorithms and LLMs, paving the way for innovative applications and interdisciplinary research in AI.

多臂赌博机大语言模型决策优化综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。