arXiv:2410.02742cs.CLcs.LG2024-10被引 2

用模拟环境训练大模型,让其学会物理推理与机器人任务

Grounding Large Language Models In Embodied Environment With Imperfect World Models

  • 用模拟器生成训练数据,结合自我优化机制提升数据质量
  • 在三个基准上使LLaMA-3性能提升1.54到2.04倍
  • 适合想提升大模型物理理解能力的研究者和开发者

尽管大型语言模型在诸多应用中表现优异,但在处理基础物理推理或执行机器人任务时仍常出错,原因在于缺乏对真实世界物理细节的直接经验。为此,我们提出一种基于不完美世界模型的大型语言模型接地方法(GLIMO),利用模拟器等代理世界模型收集并合成训练数据。GLIMO采用基于大模型代理的数据生成器,自动构建高质量、多样化的指令数据集,包含迭代式自我精炼模块以实现时间一致性体验采样、多样化的问答指令种子,以及基于检索增强生成的过往经验反思模块。综合实验表明,该方法在三个不同基准上分别将强开源模型LLaMA-3的性能提升了2.04×、1.54×和1.82×,性能可媲美甚至超越更大规模模型如GPT-4。

原文摘要 · Abstract (English)

Despite a widespread success in various applications, large language models (LLMs) often stumble when tackling basic physical reasoning or executing robotics tasks, due to a lack of direct experience with the physical nuances of the real world. To address these issues, we propose a Grounding Large language model with Imperfect world MOdel (GLIMO), which utilizes proxy world models such as simulators to collect and synthesize trining data. GLIMO incorporates an LLM agent-based data generator to automatically create high-quality and diverse instruction datasets. The generator includes an iterative self-refining module for temporally consistent experience sampling, a diverse set of question-answering instruction seeds, and a retrieval-augmented generation module for reflecting on prior experiences. Comprehensive experiments show that our approach improve the performance of strong open-source LLMs like LLaMA-3 with a performance boost of 2.04 $\times$, 1.54 $\times$, and 1.82 $\times$ across three different benchmarks, respectively. The performance is able to compete with or surpass their larger counterparts such as GPT-4.

大模型接地物理推理模拟训练机器人任务

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。