让大模型在动态环境中主动试错,发现其学习能力的瓶颈与突破点。
Can foundation models actively gather information in interactive environments to test hypotheses?
- 通过定期总结观察,激发模型自发形成元学习策略。
- 在复杂任务中,Gemini 2.5表现最优,其他模型难以持续改进。
- 揭示了模型真正短板是长期知识整合,而非即时决策。
大模型在单轮推理中表现优异,但在动态环境中的多轮探索能力不足,而这是许多现实挑战的关键。我们在‘Feature World’中测试信息收集能力,模型表现接近最优。为进一步检验多轮学习,我们采用文本版‘Alchemy’环境,该环境是元学习的基准任务,要求模型通过多次试验推断隐含因果结构。结果显示,近期大模型初始阶段无法提升性能。关键发现:定期提示模型总结观察后,涌现出元学习能力,实现跨轮次改进,并能自适应应对规则突变。尽管多数模型在简单任务中表现良好,但在Alchemy中差异显著——Gemini 2.5最佳,其次为Claude 3.7,而ChatGPT-4o和o4-mini表现不佳。这凸显Alchemy作为基准的价值。研究表明,大模型最大挑战并非即时选择有效动作,而是通过自适应策略长期整合知识。令人鼓舞的是,未来模型具备掌握这些能力的潜力。
原文摘要 · Abstract (English)
Foundation models excel at single-turn reasoning but struggle with multi-turn exploration in dynamic environments, a requirement for many real-world challenges. We evaluated these models on their ability to learn from experience, adapt, and gather information. First, in "Feature World," a simple setting for testing information gathering, models performed near-optimally. However, to test more complex, multi-trial learning, we implemented a text-based version of the "Alchemy" environment, a benchmark for meta-learning. Here, agents must deduce a latent causal structure by integrating information across many trials. In this setting, recent foundation models initially failed to improve their performance over time. Crucially, we found that prompting the models to summarize their observations at regular intervals enabled an emergent meta-learning process. This allowed them to improve across trials and even adaptively re-learn when the environment's rules changed unexpectedly. While most models handled the simple task, Alchemy revealed stark differences in robustness: Gemini 2.5 performed best, followed by Claude 3.7, while ChatGPT-4o and o4-mini struggled. This underscores Alchemy's value as a benchmark. Our findings demonstrate that the biggest challenge for foundation models is not selecting informative actions in the moment, but integrating knowledge through adaptive strategies over time. Encouragingly, there appears to be no intrinsic barrier to future models mastering these abilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。