arXiv:2409.15684cs.RO2024-09ICRA被引 1

让机器人理解人类视角,实现人机协作中的认知对齐。

SYNERGAI: Perception Alignment for Human-Robot Collaboration

  • 用3D场景图作为统一表征,驱动语言模型拆解任务并动态更新感知
  • 零样本下在ScanQA上表现媲美数据驱动模型,真实场景对齐成功率61.9%
  • 通过在线交互自动修正认知偏差,新任务成功率从3.7%提升至45.68%

近年来,大型语言模型(LLMs)在促进人机交互与协作方面展现出巨大潜力。然而,现有基于LLM的系统常忽视人与机器人之间的感知错位问题,阻碍了有效沟通与实际部署。为此,我们提出SYNERGAI——一个实现感知对齐与人机协作的统一系统。其核心采用3D场景图(3DSG)作为显式且内在的表示形式,使系统能利用LLM分解复杂任务、分配合适工具,在中间步骤中从3DSG提取信息、修改结构或生成回应。尤为重要的是,SYNERGAI包含自动机制,通过在线交互更新3DSG,实现与用户的感知错位纠正。在10个真实场景的综合实验中,该系统在零样本条件下于ScanQA任务上达到与数据驱动模型相当的性能;在对齐任务中实现61.9%的成功率,并通过知识迁移将新任务成功率从3.7%显著提升至45.68%。

原文摘要 · Abstract (English)

Recently, large language models (LLMs) have shown strong potential in facilitating human-robotic interaction and collaboration. However, existing LLM-based systems often overlook the misalignment between human and robot perceptions, which hinders their effective communication and real-world robot deployment. To address this issue, we introduce SYNERGAI, a unified system designed to achieve both perceptual alignment and human-robot collaboration. At its core, SYNERGAI employs 3D Scene Graph (3DSG) as its explicit and innate representation. This enables the system to leverage LLM to break down complex tasks and allocate appropriate tools in intermediate steps to extract relevant information from the 3DSG, modify its structure, or generate responses. Importantly, SYNERGAI incorporates an automatic mechanism that enables perceptual misalignment correction with users by updating its 3DSG with online interaction. SYNERGAI achieves comparable performance with the data-driven models in ScanQA in a zero-shot manner. Through comprehensive experiments across 10 real-world scenes, SYNERGAI demonstrates its effectiveness in establishing common ground with humans, realizing a success rate of 61.9% in alignment tasks. It also significantly improves the success rate from 3.7% to 45.68% on novel tasks by transferring the knowledge acquired during alignment.

人机协作感知对齐3D场景图LLM应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。