为机器人与人对话构建多模态共知标注体系,提升灾难救援中的协作效率。
Human-Robot Dialogue Annotation for Multi-Modal Common Ground
- 基于抽象意义表示扩展对话标注,捕捉话语语义与意图
- 设计跨轮次对话结构框架,揭示话语间关联模式
- 融合视觉信息补充理解差异,支持低带宽场景下协作
本文提出在人机对话数据上构建符号化标注,使自主系统能参与协作式自然语言对话并建立与人类伙伴的共知。在灾难救援等远程任务中,人与机器人需共同导航陌生环境,但受限于通信条件,机器人无法即时传输高质量视觉信息。对话成为有效沟通方式,辅以按需或低质量视觉信息可弥补认知差异。我们通过对话-AMR标注,捕捉单句的命题语义和言外之力;通过多轮对话结构标注,分析跨说话轮次的话语关联;初步标注并分析视觉模态如何为对话提供上下文,缓解双方对环境理解的差异。最后讨论基于这些标注实现的使用场景、架构与系统,支持物理机器人自主开展双向对话与导航。
原文摘要 · Abstract (English)
In this paper, we describe the development of symbolic representations annotated on human-robot dialogue data to make dimensions of meaning accessible to autonomous systems participating in collaborative, natural language dialogue, and to enable common ground with human partners. A particular challenge for establishing common ground arises in remote dialogue (occurring in disaster relief or search-and-rescue tasks), where a human and robot are engaged in a joint navigation and exploration task of an unfamiliar environment, but where the robot cannot immediately share high quality visual information due to limited communication constraints. Engaging in a dialogue provides an effective way to communicate, while on-demand or lower-quality visual information can be supplemented for establishing common ground. Within this paradigm, we capture propositional semantics and the illocutionary force of a single utterance within the dialogue through our Dialogue-AMR annotation, an augmentation of Abstract Meaning Representation. We then capture patterns in how different utterances within and across speaker floors relate to one another in our development of a multi-floor Dialogue Structure annotation schema. Finally, we begin to annotate and analyze the ways in which the visual modalities provide contextual information to the dialogue for overcoming disparities in the collaborators' understanding of the environment. We conclude by discussing the use-cases, architectures, and systems we have implemented from our annotations that enable physical robots to autonomously engage with humans in bi-directional dialogue and navigation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。