让大模型学会从经验中成长,提升规划能力。
Growing Through Experience: Scaling Episodic Grounding in Language Models
- 小模型经验通过新方法迁移至大模型
- 在多个任务上超越现有大模型3.45%
- 适合研究大模型经验学习与智能规划的学者
语言模型(LMs)需要强大的情景接地能力——即从过往经验中学习并应用的能力——才能在物理规划任务中表现优异。当前的情景接地方法在可扩展性和整合性方面存在瓶颈,尤其对中等规模模型(7B参数)效果有限。尽管大模型(70-405B参数)具备更优的层次化表征和丰富的预训练知识,却面临根本性的规模悖论:虽具高级抽象能力,却缺乏有效利用经验流的机制。本文提出一种可扩展的弱到强情景学习框架,能有效将小模型的经验行为迁移到大模型中。该框架结合蒙特卡洛树搜索进行结构化经验收集,并引入一种新型蒸馏方法,在保留原有语言模型能力的同时嵌入情景记忆。实验表明,该方法在多样化的规划与问答任务中,性能超越现有专有大模型3.45%。层间探查进一步显示,深层网络任务对齐显著提升,即使面对未见场景和更高复杂度规划任务,仍表现出稳定泛化能力,而基线方法则明显退化。
原文摘要 · Abstract (English)
Language models (LMs) require robust episodic grounding-the capacity to learn from and apply past experiences-to excel at physical planning tasks. Current episodic grounding approaches struggle with scalability and integration, limiting their effectiveness, especially for medium-sized LMs (7B parameters). While larger LMs (70-405B parameters) possess superior hierarchical representations and extensive pre-trained knowledge, they encounter a fundamental scale paradox: despite their advanced abstraction capabilities, they lack efficient mechanisms to leverage experience streams. We propose a scalable weak-to-strong episodic learning framework that effectively transfers episodic behaviors from smaller to larger LMs. This framework integrates Monte Carlo tree search for structured experience collection with a novel distillation method, preserving the inherent LM capabilities while embedding episodic memory. Experiments demonstrate our method surpasses state-of-the-art proprietary LMs by 3.45% across diverse planning and question-answering tasks. Layer-wise probing further indicates significant improvements in task alignment, especially within deeper LM layers, highlighting stable generalization even for previously unseen scenarios with increased planning complexity-conditions where baseline methods degrade markedly.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。