拆解大模型探索能力,发现推理强的模型更会主动探索。
Disentangling Exploration of Large Language Models by Optimal Exploitation
- 用最优收益分解法分离探索与利用,量化探索贡献。
- 多数大模型探索能力弱,且探索不足影响最终表现。
- 探索能力与推理水平正相关,适合研究提示工程影响。
在未知环境中,探索是上下文强化学习的关键能力。然而,大语言模型能否有效探索部分隐藏的状态空间仍不明确。本文将探索设为唯一目标,让智能体专注于收集能提升未来回报的信息。在此框架下,我们指出仅测量智能体回报不足以公平评估探索能力,因此基于最优可实现回报,将缺失奖励分解为探索与利用两部分。不同模型的实验表明,多数模型探索能力薄弱,且探索不足。但发现探索表现与推理能力呈正相关。该分解方法可揭示提示工程带来的行为差异,为优化探索任务性能提供有力工具。
原文摘要 · Abstract (English)
Exploration is a crucial skill for in-context reinforcement learning in unknown environments. However, it remains unclear if large language models can effectively explore a partially hidden state space. This work isolates exploration as the sole objective, tasking an agent with gathering information that enhances future returns. Within this framework, we argue that measuring agent returns is not sufficient for a fair evaluation. Hence, we decompose missing rewards into their exploration and exploitation components based on the optimal achievable return. Experiments with various models reveal that most struggle to explore the state space, and weak exploration is insufficient. Nevertheless, we found a positive correlation between exploration performance and reasoning capabilities. Our decomposition can provide insights into differences in behaviors driven by prompt engineering, offering a valuable tool for refining performance in exploratory tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。