arXiv:2503.16492cs.HCcs.RO2025-03中稿 · publication in IEE…被引 16

用眼神+语音让机器人更懂人,特别适合行动不便者

FAM-HRI: Foundation-Model Assisted Multi-Modal Human-Robot Interaction Combining Gaze and Speech

  • 结合语言与注视信号,用大模型理解用户意图
  • 实测任务成功率高,交互时间短于传统方式
  • 轻量级设备+开源代码,适合残障人士辅助使用

高效的人机交互(HRI)对提升实际机器人应用的可及性与可用性至关重要。现有方案多依赖手势或语言单一指令,导致交互低效且模糊,尤其对肢体障碍用户不友好。本文提出FAM-HRI,一种基于基础模型的多模态人机交互框架,融合语言与注视输入。通过轻量级Meta ARIA眼镜实时捕捉多模态信号,并利用大语言模型(LLMs)融合用户意图与场景上下文,实现直观精准的机器人操控。方法有效识别注视固定时间区间,降低注视动态带来的噪声。实验表明,该系统在任务执行中达到高成功率,同时保持低交互时长,为运动能力受限者提供实用解决方案。相关系统设计、算法与实现已开源:https://github.com/laiyuzhi/FAM-HRI。

原文摘要 · Abstract (English)

ffective Human-Robot Interaction (HRI) is crucial for enhancing accessibility and usability in real-world robotics applications. However, existing solutions often rely on gesture- only or language-only commands, making interaction inefficient and ambiguous, particularly for users with physical impairments. In this paper, we introduce FAM-HRI, an efficient multimodal framework for HRI that integrates language and gaze inputs via foundation models. By leveraging lightweight Meta ARIA glasses, our system captures real-time multimodal signals and utilizes large language models (LLMs) to fuse user intention with scene context, enabling intuitive and precise robot manipulation. Our method accurately determines the gaze fixation time interval, reducing noise caused by the gaze dynamic nature. Experimental evaluations demonstrate that FAM-HRI achieves a high success rate in task execution while maintaining a low interaction time, providing a practical solution for individuals with limited physical mobility or motor impairments. To support the community, we have released our system design, algorithms, and solutions at https://github.com/laiyuzhi/FAM-HRI.

人机交互多模态视觉语言模型无障碍设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。