LLM做决策时,反馈反而让复杂任务表现变差。
Feedback-Induced Performance Decline in LLM-Based Decision-Making
- 用结构化提示让LLM在线决策,比传统强化学习快
- 简单任务中LLM表现好,复杂场景下性能下降
- 反馈机制会引发混淆,适合研究混合策略的学者
大型语言模型(LLMs)从自然语言问题描述中提取上下文的能力,引发了其在自主决策场景中的适用性问题。本文研究了这些模型在马尔可夫决策过程(MDPs)中的行为。尽管传统强化学习(RL)依赖迭代探索,而预训练的LLMs能利用先验知识实现快速适应。我们对比了基于LLM的零样本方法与经典RL方法在序列决策任务中的表现。结果表明,虽然在简单环境中LLM初始表现更优,但在复杂场景中缺乏微调或额外引导时,其规划与推理能力受限。更重要的是,本研究发现反馈机制本意提升决策质量,却常导致混乱,使复杂环境下的性能下降。这凸显了未来需探索混合策略、微调和先进记忆集成以增强LLM决策能力。
原文摘要 · Abstract (English)
The ability of Large Language Models (LLMs) to extract context from natural language problem descriptions naturally raises questions about their suitability in autonomous decision-making settings. This paper studies the behaviour of these models within a Markov Decision Process (MDPs). While traditional reinforcement learning (RL) strategies commonly employed in this setting rely on iterative exploration, LLMs, pre-trained on diverse datasets, offer the capability to leverage prior knowledge for faster adaptation. We investigate online structured prompting strategies in sequential decision making tasks, comparing the zero-shot performance of LLM-based approaches to that of classical RL methods. Our findings reveal that although LLMs demonstrate improved initial performance in simpler environments, they struggle with planning and reasoning in complex scenarios without fine-tuning or additional guidance. Our results show that feedback mechanisms, intended to improve decision-making, often introduce confusion, leading to diminished performance in intricate environments. These insights underscore the need for further exploration into hybrid strategies, fine-tuning, and advanced memory integration to enhance LLM-based decision-making capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。