arXiv:2410.16804cs.RO2024-10中稿 · IEEE/RSJ Internati…被引 6

用大模型+知识图谱让机器人听懂模糊指令,减少用户反复解释。

Combining Ontological Knowledge and Large Language Model for User-Friendly Service Robots

  • 融合大语言模型与知识图谱,理解模糊任务指令
  • 在“拿东西”任务中降低30%用户交互次数
  • 适合需要自然对话的智能服务机器人场景

生活方式支持型机器人正成为重要研究方向,期望其能承担清洁、摆桌、取物等家务。随着大语言模型(LLMs)和视觉语言模型(VLMs)的发展,机器人交互能力显著提升。本文聚焦于‘拿东西’类任务,即机器人根据用户模糊指令取物。此前方法通过扩展知识图谱处理环境信息以解析歧义,但遇到无法解决的模糊时仍需用户澄清。本文改进方案,引入大语言模型提供常识性知识,与知识图谱协同工作,有效抑制幻觉,减少用户干预需求,提升系统可用性。我们构建了融合双知识源的系统,并在‘拿东西’任务上验证其有效性,旨在实现更流畅高效的机器人辅助体验。

原文摘要 · Abstract (English)

Lifestyle support through robotics is an increasingly promising field, with expectations for robots to take over or assist with chores like floor cleaning, table setting and clearing, and fetching items. The growth of AI, particularly foundation models, such as large language models (LLMs) and visual language models (VLMs), is significantly shaping this sector. LLMs, by facilitating natural interactions and providing vast general knowledge, are proving invaluable for robotic tasks. This paper zeroes in on the benefits of LLMs for "bring-me" tasks, where robots fetch specific items for users, often based on vague instructions. Our previous efforts utilized an ontology extended to handle environmental data to decipher such vagueness, but faced limitations when unresolvable ambiguities required user intervention for clarity. Here, we enhance our approach by integrating LLMs for providing additional commonsense knowledge, pairing it with ontological data to mitigate the issue of hallucinations and reduce the need for user queries, thus improving system usability. We present a system that merges these knowledge bases and assess its efficacy on "bring-me" tasks, aiming to provide a more seamless and efficient robotic assistance experience.

服务机器人大模型知识图谱自然交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。