智能机器人融合跟随、对话与跌倒监测,提升养老院照护体验。
Multimodal-Language-Model-Driven Interaction and Companionship for Service Robots in Elderly-Care Facilities

- 用主动云台+视觉追踪实现稳定跟随,遮挡时也不丢失目标。
- 大语言模型理解语音指令,支持导航与任务执行,交互自然流畅。
- 视觉大模型实时分析姿态,精准识别跌倒和异常姿势并报警。
服务机器人正被广泛部署于养老机构以减轻护理人员负担、提升日常照护质量。然而,现有研究多聚焦单一功能,缺乏持续陪伴、自然交互与安全监控的整合能力。本文提出一种智能陪伴机器人系统,融合主动视觉人形追踪、基于大语言模型的实时语音交互以理解意图并执行任务,以及基于视觉语言模型的安全监控,用于跌倒检测与异常体位评估。感知层通过主动云台确保在遮挡或突发移动时持续追踪用户;交互层利用大语言模型解析语音请求并映射为机器人动作,实现引导与语义导航;同时,视觉语言模型安全代理持续分析视觉输入,检测跌倒或异常姿态,并在必要时触发紧急响应。实验结果表明,该系统能可靠地跟随与交互人类,并有效检测潜在跌倒风险,保障用户安全。
原文摘要 · Abstract (English)
Service robots are increasingly deployed in elderly-care facilities to alleviate caregiver workload and enhance the quality of daily care. However, most existing studies focus on isolated service functions and lack integrated capabilities for continuous companionship, natural interaction, and safety monitoring. In this paper, we present an intelligent companion robot system that unifies active visual human-following, real-time LLM-driven speech interaction for intent understanding and task execution, and VLM-based safety monitoring for fall detection and abnormal posture assessment. The perception layer ensures robust human tracking and uses an active gimbal to maintain the user in view during occlusions or abrupt movements. At the interaction layer, a Large Language Model interprets spoken requests and maps them to robot actions, enabling escorting and semantic navigation. Simultaneously, a VLM-based safety agent continuously analyzes visual observations to detect fall-related or abnormal postures and triggers emergency responses when necessary. Experimental results demonstrate the system's ability to reliably follow and interact with humans, while effectively detecting potential falls to ensure user safety.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。