arXiv:2502.11733cs.CL2025-02中稿 · The 28th Internati…被引 1

测试大模型在真实场景中边学边做的能力,发现闭源模型表现差、新词理解难。

Plant in Cupboard, Orange on Rably, Inat Aphone. Benchmarking Incremental Learning of Situation and Language Model using a Text-Simulated Situated Environment

  • 用文本模拟家居环境,让模型通过互动逐步学习物品位置和任务
  • 闭源大模型比开源模型差很多,对虚构单词几乎无法理解
  • 适合研究智能体学习、语言模型常识推理的学者参考

大型语言模型(LLMs)不仅是聊天机器人,更是智能体系统中的核心组件,其常识知识直接影响其作为基于语言的规划器在具身或情境化行动中的表现。我们通过一个基于文本的环境,评估了LLMs的增量学习(基于环境反馈)与受控的上下文学习能力。设计了三项挑战性实验:一是在典型房间中通过交互逐步发现日常物品并完成任务;二是提供简短的位置信息,测试上下文学习的速度与效率;三是引入合成伪英语词汇,考察模型如何从环境反馈中推断未知词义。结果显示,更大的商业模型与开源模型之间存在显著性能差距,但几乎所有模型在合成词汇任务中均表现不佳。

原文摘要 · Abstract (English)

Large Language Models (LLMs) serve not only as chatbots but as key components in agent systems, where their common-sense knowledge significantly impacts performance as language-based planners for situated or embodied action. We assess LLMs' incremental learning (based on feedback from the environment), and controlled in-context learning abilities using a text-based environment. We introduce challenging yet interesting set of experiments to test i) how agents can incrementally solve tasks related to every day objects in typical rooms in a house where each of them are discovered by interacting within the environment, ii) controlled in-context learning abilities and efficiency of agents by providing short info about locations of objects and rooms to check how faster the task can be solved, and finally iii) using synthetic pseudo-English words to gauge how well LLMs are at inferring meaning of unknown words from environmental feedback. Results show that larger commercial models have a substantial gap in performance compared to open-weight but almost all models struggle with the synthetic words experiments.

语言模型增量学习常识推理智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。