arXiv:2507.09617cs.AIcs.RO2025-07

用多模态大模型+知识图谱让机器人跨平台理解与执行任务

Bridging Bots: from Perception to Action via Multimodal-LMs and Knowledge Graphs

  • 融合多模态大模型感知能力与知识图谱结构化知识
  • GPT-o1和LLaMA 4 Maverick生成的知识图谱一致性最优
  • 适合需要跨平台兼容的智能服务机器人研发者

个人服务机器人被部署于家庭环境,以支持老年人及需要协助人群的日常生活。这些机器人需感知复杂动态环境、理解任务并执行合适行为。然而现有系统依赖专有且硬编码的方案,绑定特定软硬件,导致系统孤岛化,难以跨平台适配与扩展。本研究提出一种神经符号框架,结合多模态语言模型(MLMs)的感知能力与知识图谱(KGs)及本体的结构化表示,实现机器人应用的互操作性。该框架生成符合本体规范的KGs,可独立于平台指导机器人行为。我们集成五种多模态模型(三个LLaMA和两个GPT模型),通过不同神经符号交互模式进行评估。结果表明,GPT-o1和LLaMA 4 Maverick在多轮运行中表现最优;但新模型未必更优,凸显集成策略在生成符合本体的KGs中的关键作用。

原文摘要 · Abstract (English)

Personal service robots are deployed to support daily living in domestic environments, particularly for elderly and individuals requiring assistance. These robots must perceive complex and dynamic surroundings, understand tasks, and execute context-appropriate actions. However, current systems rely on proprietary, hard-coded solutions tied to specific hardware and software, resulting in siloed implementations that are difficult to adapt and scale across platforms. Ontologies and Knowledge Graphs (KGs) offer a solution to enable interoperability across systems, through structured and standardized representations of knowledge and reasoning. However, symbolic systems such as KGs and ontologies struggle with raw and noisy sensory input. In contrast, multimodal language models are well suited for interpreting input such as images and natural language, but often lack transparency, consistency, and knowledge grounding. In this work, we propose a neurosymbolic framework that combines the perceptual strengths of multimodal language models with the structured representations provided by KGs and ontologies, with the aim of supporting interoperability in robotic applications. Our approach generates ontology-compliant KGs that can inform robot behavior in a platform-independent manner. We evaluated this framework by integrating robot perception data, ontologies, and five multimodal models (three LLaMA and two GPT models), using different modes of neural-symbolic interaction. We assess the consistency and effectiveness of the generated KGs across multiple runs and configurations, and perform statistical analyzes to evaluate performance. Results show that GPT-o1 and LLaMA 4 Maverick consistently outperform other models. However, our findings also indicate that newer models do not guarantee better results, highlighting the critical role of the integration strategy in generating ontology-compliant KGs.

机器人多模态知识图谱神经符号

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。