arXiv:2409.17702cs.ROcs.AI2024-09被引 10

用分层结构让机器人回忆生活经历并回答问题

Episodic Memory Verbalization using Hierarchical Representations of Life-Long Robot Experience

  • 构建树状记忆结构,从原始数据到自然语言事件逐级抽象
  • 仅需少量示例即可处理数月机器人经验,支持高效问答
  • 适合长期服役的家用机器人或人机交互研究者

机器人经验的言语化,即对过往行为进行摘要与问答,是提升人机交互的关键。以往方法依赖规则系统或微调深度模型,仅能处理几分钟的短时数据流,泛化能力差。本文利用大规模预训练模型,在零样本或少样本条件下,针对长期机器人经验的言语化任务。我们从机器人经验流中构建树状数据结构:底层为原始感知与本体感觉数据,高层抽象为自然语言事件概念。基于此分层表示,使用大语言模型作为代理,根据用户提问动态展开(初始压缩的)树节点,搜索相关记忆。该方法在扩展至数月机器人数据时仍保持低计算开销。我们在模拟家庭机器人数据、人类第一视角视频及真实机器人记录上验证,展示了方法的灵活性与可扩展性。

原文摘要 · Abstract (English)

Verbalization of robot experience, i.e., summarization of and question answering about a robot's past, is a crucial ability for improving human-robot interaction. Previous works applied rule-based systems or fine-tuned deep models to verbalize short (several-minute-long) streams of episodic data, limiting generalization and transferability. In our work, we apply large pretrained models to tackle this task with zero or few examples, and specifically focus on verbalizing life-long experiences. For this, we derive a tree-like data structure from episodic memory (EM), with lower levels representing raw perception and proprioception data, and higher levels abstracting events to natural language concepts. Given such a hierarchical representation built from the experience stream, we apply a large language model as an agent to interactively search the EM given a user's query, dynamically expanding (initially collapsed) tree nodes to find the relevant information. The approach keeps computational costs low even when scaling to months of robot experience data. We evaluate our method on simulated household robot data, human egocentric videos, and real-world robot recordings, demonstrating its flexibility and scalability.

机器人记忆语言模型长期学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。