用实时多模态大模型和工具调用实现机器人情境对话中的动态感知。
A Modern System Recipe for Situated Embodied Human-Robot Conversation with Real-Time Multimodal LLMs and Tool-Calling
- 用小工具集控制机器人注意力与主动感知
- 六种家庭场景下工具决策准确率达85%以上
- 适合做具身智能交互系统的快速原型开发
情境化具身对话要求机器人在严格延迟约束下,将实时对话与主动感知交织进行:决定看什么、何时看、说什么。本文提出一种简洁、轻量的系统方案,将实时多模态语言模型与一组小型工具接口结合,用于注意力控制和主动感知。我们评估了四种系统变体在六种家庭场景下的表现,这些场景需频繁切换注意力并逐步扩大感知范围。通过与人工标注对比,评估每轮对话中工具决策的正确性,并收集用户对交互质量的主观评分。结果表明,实时多模态大语言模型结合工具调用,在实际情境化具身对话中具有显著潜力。
原文摘要 · Abstract (English)
Situated embodied conversation requires robots to interleave real-time dialogue with active perception: deciding what to look at, when to look, and what to say under tight latency constraints. We present a simple, minimal system recipe that pairs a real-time multimodal language model with a small set of tool interfaces for attention and active perception. We study six home-style scenarios that require frequent attention shifts and increasing perceptual scope. Across four system variants, we evaluate turn-level tool-decision correctness against human annotations and collect subjective ratings of interaction quality. Results indicate that real-time multimodal large language models and tool use for active perception is a promising direction for practical situated embodied conversation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。