语言模型在有限交互下探索能力差,且越给预算表现越弱。
Failing to Explore: Language Models on Interactive Tasks
- 设计三类可控难度的交互任务,测试模型探索能力
- 多数模型探索不足,性能远低于简单启发式方法
- 并行执行预算和历史摘要能有效提升探索效果
我们评估语言模型在有限交互预算下探索交互环境的能力。引入三个参数化任务,涵盖连续与离散环境,探索难度可调。在主流模型中,普遍发现系统性探索不足与次优解,性能常显著低于简单的探索-利用启发式基线,且随着预算增加,表现提升微弱。最后研究两种轻量级干预:将固定预算拆分为并行执行,尽管理论上无增益,但意外提升性能;定期总结交互历史,有助于保留关键发现并进一步改善探索。
原文摘要 · Abstract (English)
We evaluate language models on their ability to explore interactive environments under a limited interaction budget. We introduce three parametric tasks with controllable exploration difficulty, spanning continuous and discrete environments. Across state-of-the-art models, we find systematic under-exploration and suboptimal solutions, with performance often significantly worse than simple explore--exploit heuristic baselines and scaling weakly as the budget increases. Finally, we study two lightweight interventions: splitting a fixed budget into parallel executions, which surprisingly improves performance despite a no-gain theoretical result for our tasks, and periodically summarizing the interaction history, which preserves key discoveries and further improves exploration.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。