arXiv:2411.13410cs.LG2024-11综述被引 16

用人类和大模型反馈提升强化学习在复杂环境中的表现

A Survey On Enhancing Reinforcement Learning in Complex Environments: Insights from Human and LLM Feedback

  • 结合人类或大模型的自然语言反馈指导强化学习
  • 显著缓解高维观测空间下的样本低效问题
  • 适合关注智能体协作与高效学习的研究者

强化学习(RL)是机器学习的活跃领域,具有解决现实挑战的巨大潜力。然而,面对观测空间庞大的复杂环境时,现有方法常因样本效率低下和学习周期过长而表现不佳,这被称为维度诅咒。该问题使决策变得复杂,需在注意力与决策间取得平衡。通过引入人类或大型语言模型(LLMs)的反馈,强化学习代理可增强鲁棒性与适应性,实现更优性能和更快学习。此类反馈以自然语言等多模态、多层次形式传递,帮助智能体识别关键环境线索并优化决策过程。本文综述聚焦两大方向:一是探讨人类或大模型如何与强化学习代理协作,促进最优行为并加速学习;二是系统梳理针对高维观测空间环境复杂性的研究进展。

原文摘要 · Abstract (English)

Reinforcement learning (RL) is one of the active fields in machine learning, demonstrating remarkable potential in tackling real-world challenges. Despite its promising prospects, this methodology has encountered with issues and challenges, hindering it from achieving the best performance. In particular, these approaches lack decent performance when navigating environments and solving tasks with large observation space, often resulting in sample-inefficiency and prolonged learning times. This issue, commonly referred to as the curse of dimensionality, complicates decision-making for RL agents, necessitating a careful balance between attention and decision-making. RL agents, when augmented with human or large language models' (LLMs) feedback, may exhibit resilience and adaptability, leading to enhanced performance and accelerated learning. Such feedback, conveyed through various modalities or granularities including natural language, serves as a guide for RL agents, aiding them in discerning relevant environmental cues and optimizing decision-making processes. In this survey paper, we mainly focus on problems of two-folds: firstly, we focus on humans or an LLMs assistance, investigating the ways in which these entities may collaborate with the RL agent in order to foster optimal behavior and expedite learning; secondly, we delve into the research papers dedicated to addressing the intricacies of environments characterized by large observation space.

强化学习人类反馈大模型复杂环境

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。