构建了用于人机协作探索的多模态对话语料库,支持机器人理解人类指令。
SCOUT: A Situated and Multi-Modal Human-Robot Dialogue Corpus
- 通过真人远程操控机器人收集真实对话数据,涵盖语音与视觉信息。
- 包含278段对话、8.9万条语句和31万词,每段平均320条,配以图像与地图数据。
- 标注了语义意图与对话结构,适合研究人机交互中自然语言理解与导航任务。
我们提出了情境化对话理解交易语料库(SCOUT),一个聚焦于协作探索任务的人-机对话多模态数据集。该语料库基于多个‘巫师在幕后’实验构建,参与者通过口头指令远程控制机器人移动并获取环境信息。共收集278段对话,包含89,056条语句和310,095个词,平均每段对话320条。对话数据与实验期间采集的5,785张图像及30张地图对齐。语料库已标注抽象语义表示(AMR)和对话-AMR,以识别说话者意图;同时标注事务单元与关系,用于揭示对话结构模式。这些数据已用于开发自主人机系统,并推动对人类如何向机器人表达的开放问题研究。本研究公开该语料库,旨在加速自主、情境化人机对话的发展,特别是在需发现环境细节的导航任务中。
原文摘要 · Abstract (English)
We introduce the Situated Corpus Of Understanding Transactions (SCOUT), a multi-modal collection of human-robot dialogue in the task domain of collaborative exploration. The corpus was constructed from multiple Wizard-of-Oz experiments where human participants gave verbal instructions to a remotely-located robot to move and gather information about its surroundings. SCOUT contains 89,056 utterances and 310,095 words from 278 dialogues averaging 320 utterances per dialogue. The dialogues are aligned with the multi-modal data streams available during the experiments: 5,785 images and 30 maps. The corpus has been annotated with Abstract Meaning Representation and Dialogue-AMR to identify the speaker's intent and meaning within an utterance, and with Transactional Units and Relations to track relationships between utterances to reveal patterns of the Dialogue Structure. We describe how the corpus and its annotations have been used to develop autonomous human-robot systems and enable research in open questions of how humans speak to robots. We release this corpus to accelerate progress in autonomous, situated, human-robot dialogue, especially in the context of navigation tasks where details about the environment need to be discovered.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。