arXiv:2412.01441cs.AIcs.LG2024-12ICML被引 38

测试大模型在超长上下文下从大量示范中学习决策的能力

LMAct: A Benchmark for In-Context Imitation Learning with Long Multimodal Demonstrations

  • 构建涵盖百万词元的长上下文基准,评估模型从多示范中学习
  • 512个完整示范对多数任务提升有限,仅少数模型持续进步
  • 适合研究长序列理解、具身智能与示范学习的学者使用

本文提出一个基准,用于检验当前前沿模型在超长上下文(最高达一百万词元)下进行多模态决策的能力,并探究模型能否通过上下文中的大量专家示范实现学习。我们在一系列简单交互式决策任务上评估了Claude 3.5 Sonnet、Gemini 1.5 Flash、Gemini 1.5 Pro、Gemini 2.0 Flash Experimental、GPT-4o、o1-mini、o1-preview和o1的表现,包括井字棋、国际象棋、Atari游戏、网格世界导航、填字游戏和模拟猎豹控制。实验考察了上下文中示范数量从零到512个完整回合的变化影响。结果显示,多数任务中模型难以达到专家水平,增加示范数量也常无显著效果;仅有少数模型在部分任务上随示范增多而持续提升。我们还研究了将观测编码为文本或图像的影响,以及思维链提示的作用。为量化其他方法的效果并推动未来创新,我们开源了该涵盖零样本、少样本和多样本场景的统一评估基准。

原文摘要 · Abstract (English)

In this paper, we present a benchmark to pressure-test today's frontier models' multimodal decision-making capabilities in the very long-context regime (up to one million tokens) and investigate whether these models can learn from large numbers of expert demonstrations in their context. We evaluate the performance of Claude 3.5 Sonnet, Gemini 1.5 Flash, Gemini 1.5 Pro, Gemini 2.0 Flash Experimental, GPT-4o, o1-mini, o1-preview, and o1 as policies across a battery of simple interactive decision-making tasks: playing tic-tac-toe, chess, and Atari, navigating grid worlds, solving crosswords, and controlling a simulated cheetah. We study increasing amounts of expert demonstrations in the context $\unicode{x2013}$ from no demonstrations to 512 full episodes. Across our tasks, models rarely manage to fully reach expert performance, and often, presenting more demonstrations has little effect. Some models steadily improve with more demonstrations on a few tasks. We investigate the effect of encoding observations as text or images and the impact of chain-of-thought prompting. To help quantify the impact of other approaches and future innovations, we open source our benchmark that covers the zero-, few-, and many-shot regimes in a unified evaluation.

示范学习长上下文多模态基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。