arXiv:2409.16900cs.ROcs.AI2024-09中稿 · Version of a confe…被引 5

让大模型通过身体与社会互动实现真正理解语言。

A Roadmap for Embodied and Social Grounding in LLMs

  • 以身体为参照,构建感知-行动闭环
  • 强调时间连贯性,形成自我关联的体验
  • 需具备社交能力才能共享语义共识

大型语言模型(LLMs)与机器人系统的融合正推动机器人领域变革,不仅提升沟通能力,还在多模态输入处理、高层推理和规划生成方面展现强大潜力。将LLMs的知识与现实世界进行具身化连接,是发挥其在机器人中效率的关键路径。然而,仅通过多模态方式或机器人本体与外部世界建立联系,并不足以使模型真正理解其所操作的语言。借鉴人类认知机制,本文提出实现语言理解必须具备三个要素:以主动具身系统为感知环境的参考点;具有时间结构的体验以支持与外部世界的一致性、自我相关交互;以及社会技能,以建立共享的共同经验基础。由此构建出面向大模型具身与社会性接地的路线图。

原文摘要 · Abstract (English)

The fusion of Large Language Models (LLMs) and robotic systems has led to a transformative paradigm in the robotic field, offering unparalleled capabilities not only in the communication domain but also in skills like multimodal input handling, high-level reasoning, and plan generation. The grounding of LLMs knowledge into the empirical world has been considered a crucial pathway to exploit the efficiency of LLMs in robotics. Nevertheless, connecting LLMs' representations to the external world with multimodal approaches or with robots' bodies is not enough to let them understand the meaning of the language they are manipulating. Taking inspiration from humans, this work draws attention to three necessary elements for an agent to grasp and experience the world. The roadmap for LLMs grounding is envisaged in an active bodily system as the reference point for experiencing the environment, a temporally structured experience for a coherent, self-related interaction with the external world, and social skills to acquire a common-grounded shared experience.

具身智能语言理解机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。