用凝视+语音实现更精准的机器人协作,适合行动不便者使用
Glance-Say: Multimodal Human-Robot Collaboration and Intent Recognition via Sticky Glance
- 通过融合几何距离与方向信息稳定凝视信号,提升目标识别鲁棒性
- 静态目标选择准确率达97%,动态目标追踪率92%,任务耗时减少
- 支持连续人机协同控制,适合残障人士在多物体环境中高效操作
凝视与语音是行动障碍人群理想的交互方式,但在多物体环境中,因微跳动、语义模糊和视角变化,意图识别仍具挑战。本文提出一种辅助机器人操作的多模态交互框架。我们设计了粘性凝视算法,联合累积几何距离与方向证据,实现稳定实时的目标选择与切换。进一步提出Glance-Say交互范式,凝视指定对象、语音指定动作,并采用连续共享控制机制,提供高响应度机器人运动与人机反馈。实验表明,动态目标追踪率为0.92,静态目标选择准确率为0.97,任务时长显著缩短。结果表明,该方法在鲁棒性、效率与可用性方面优于现有范式。
原文摘要 · Abstract (English)
Gaze and speech are promising interaction modalities for individuals with motor impairments, yet robust intent recognition in multi-object environments remains challenging due to micro-saccades, semantic ambiguity, and viewpoint changes. This paper presents a multimodal interaction framework for assistive robotic manipulation. We propose a sticky-glance algorithm that stabilizes gaze-based intent by jointly accumulating geometric distance and directional evidence, enabling robust real-time target selection and switching. We further introduce Glance-Say, a gaze-speech interaction paradigm in which gaze specifies objects and speech specifies actions, together with a continuous shared-control scheme that provides high-readiness robot motion and human-in-the-loop feedback. Experiments demonstrate a tracking rate of 0.92 for moving targets, selection accuracy of 0.97 for static targets, and reduced task duration. These results indicate improved robustness, efficiency, and usability over representative interaction paradigms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。