用AI生成动态视觉辅助,让人类指挥机器人更直观高效。
GenComUI: Exploring Generative Visual Aids as Medium to Support Task-Oriented Human-Robot Communication
- 基于大模型实时生成地图标注、路径指示等视觉辅助
- 用户实验显示视觉反馈显著提升沟通效率与准确性
- 适合需要复杂指令交互的机器人协作场景
本研究探索在人机任务通信中融合生成式视觉辅助的效果。我们开发了GenComUI系统,该系统依托大语言模型,动态生成地图标注、路径指示和动画等上下文视觉辅助,以支持口头任务指令,并帮助生成定制化机器人任务程序。该系统基于一项形成性研究,考察了人类在空间任务中如何使用外部视觉工具辅助口头交流。通过20名用户的对照实验(对比仅语音基线),结果表明:生成式视觉辅助通过定性和定量分析均能提升口头任务沟通效果,提供持续视觉反馈,促进自然高效的人机互动。研究还提炼出一系列设计启示,强调动态生成视觉辅助在人机交互中的有效媒介作用。这些发现凸显了生成式视觉辅助在复杂人机通信及基于大模型的用户端开发场景中的应用潜力。
原文摘要 · Abstract (English)
This work investigates the integration of generative visual aids in human-robot task communication. We developed GenComUI, a system powered by large language models that dynamically generates contextual visual aids (such as map annotations, path indicators, and animations) to support verbal task communication and facilitate the generation of customized task programs for the robot. This system was informed by a formative study that examined how humans use external visual tools to assist verbal communication in spatial tasks. To evaluate its effectiveness, we conducted a user experiment (n = 20) comparing GenComUI with a voice-only baseline. The results demonstrate that generative visual aids, through both qualitative and quantitative analysis, enhance verbal task communication by providing continuous visual feedback, thus promoting natural and effective human-robot communication. Additionally, the study offers a set of design implications, emphasizing how dynamically generated visual aids can serve as an effective communication medium in human-robot interaction. These findings underscore the potential of generative visual aids to inform the design of more intuitive and effective human-robot communication, particularly for complex communication scenarios in human-robot interaction and LLM-based end-user development.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。