用大模型生成机器人协作视觉信号,显著提升沟通效率与用户体验。
SiSCo: Signal Synthesis for Effective Human-Robot Communication Via Large Language Models
- 结合大模型与混合现实技术,自动生成上下文感知的视觉信号。
- 任务完成时间减少73%,成功率提升18%,认知负荷降低46%。
- 适合需要快速生成协作信号的机器人应用与人机交互研究者。
高效的人机协作依赖于可靠的通信渠道,其中视觉信号因其直观性而尤为重要。然而,设计直观的视觉提示通常需要大量资源和专业知识。大型语言模型(LLMs)为提升人机交互并革新情境感知视觉提示的生成方式提供了新路径。为此,我们提出SiSCo——一种将大模型计算能力与混合现实技术相结合的新框架,用于简化人机协作中视觉提示的创建。结果表明,相比基线自然语言信号,SiSCo将团队协作任务的完成时间减少了约73%,任务成功率提升了18%。此外,参与者认知负荷降低了46%(基于NASA-TLX量表),且对未见过物体生成的即时信号评价高于平均水平。为促进后续发展与社区参与,我们已在GitHub公开了SiSCo的完整实现及相关材料。
原文摘要 · Abstract (English)
Effective human-robot collaboration hinges on robust communication channels, with visual signaling playing a pivotal role due to its intuitive appeal. Yet, the creation of visually intuitive cues often demands extensive resources and specialized knowledge. The emergence of Large Language Models (LLMs) offers promising avenues for enhancing human-robot interactions and revolutionizing the way we generate context-aware visual cues. To this end, we introduce SiSCo--a novel framework that combines the computational power of LLMs with mixed-reality technologies to streamline the creation of visual cues for human-robot collaboration. Our results show that SiSCo improves the efficiency of communication in human-robot teaming tasks, reducing task completion time by approximately 73% and increasing task success rates by 18% compared to baseline natural language signals. Additionally, SiSCo reduces cognitive load for participants by 46%, as measured by the NASA-TLX subscale, and receives above-average user ratings for on-the-fly signals generated for unseen objects. To encourage further development and broader community engagement, we provide full access to SiSCo's implementation and related materials on our GitHub repository.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。