arXiv:2410.14141cs.ROcs.CL2024-10被引 7

让机器人在危险场景下更懂人话,对话更安全有效。

Coherence-Driven Multimodal Safety Dialogue with Active Learning for Embodied Agents

  • 用对话连贯性增强机器人对安全场景的理解能力。
  • 在1000个真实危险场景上验证,对话安全性与用户满意度显著提升。
  • 适合需要实时安全交互的具身机器人应用。

在协助日常任务时,机器人需准确理解视觉线索并在多样安全关键情境中有效回应,如地面上的尖锐物品。为此,我们提出M-CoDAL多模态对话系统,专为具身智能体设计,以增强其在安全关键情境中的理解和沟通能力。该系统利用话语连贯关系提升上下文理解。为训练系统,我们引入一种基于聚类的主动学习机制,通过外部大语言模型识别有信息量的样本实例。评估基于新构建的多模态数据集,包含从2000张Reddit图像中提取的1000个安全违规事件,由大模态模型(LMM)标注并经人工验证。结果表明,该方法显著提升安全情境解决率、用户情感体验及对话安全性。进一步,我们将系统部署于Hello Robot Stretch机器人,并开展受试者内用户研究,参与者在不同严重程度的安全场景中与机器人互动,接收本模型和基于OpenAI ChatGPT的基线系统的干预。研究结果证实并扩展了自动评估发现:本系统在真实具身场景中更具说服力。

原文摘要 · Abstract (English)

When assisting people in daily tasks, robots need to accurately interpret visual cues and respond effectively in diverse safety-critical situations, such as sharp objects on the floor. In this context, we present M-CoDAL, a multimodal-dialogue system specifically designed for embodied agents to better understand and communicate in safety-critical situations. The system leverages discourse coherence relations to enhance its contextual understanding and communication abilities. To train this system, we introduce a novel clustering-based active learning mechanism that utilizes an external Large Language Model (LLM) to identify informative instances. Our approach is evaluated using a newly created multimodal dataset comprising 1K safety violations extracted from 2K Reddit images. These violations are annotated using a Large Multimodal Model (LMM) and verified by human annotators. Results with this dataset demonstrate that our approach improves resolution of safety situations, user sentiment, as well as safety of the conversation. Next, we deploy our dialogue system on a Hello Robot Stretch robot and conduct a within-subject user study with real-world participants. In the study, participants role-play two safety scenarios with different levels of severity with the robot and receive interventions from our model and a baseline system powered by OpenAI's ChatGPT. The study results corroborate and extend the findings from the automated evaluation, showing that our proposed system is more persuasive in a real-world embodied agent setting.

具身智能安全对话主动学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。