通过音视频融合实现机器人自主识别用户互动意图
Initiation of Interaction Detection Framework using a Nonverbal Cue for Human-Robot Interaction

- 结合声音定位与人脸追踪,判断用户是否面向机器人
- 无需语音关键词,仅凭注视或说话即可触发互动检测
- 适用于家庭场景的移动机器人,适合人机交互研究者
本文提出一种基于音视频传感器融合的无关键词人机交互启动检测框架,应用于家庭环境。机器人内置音频与视觉传感器,可配合外部视觉传感器实现稳定的人体检测与跟踪。当用户面对机器人说话时,系统通过声源定位与人体追踪信息确定其位置,并检测其面部是否朝向机器人以判定互动启动。若用户未直接说话,但持续注视机器人超过预设时间,系统同样可识别为互动启动。设计并验证了状态转移模型,所有组件均在机器人操作系统(ROS)环境中实现与集成。
原文摘要 · Abstract (English)
This paper describes an initiation of interaction(IoI) detection framework without keywords for human-robot interaction(HRI) based on audio and vision sensor fusion in a domestic environment. In the proposed framework, the robot has its own audio and vision sensors, and can employ external vision sensor for stable human detection and tracking. When the user starts to speak while looking at the robot, the robot can localize his or her position by its sound source localization together with human tracking information. Then the robot can detect the IoI if it perceives the face of the speaker faces the robot. In case that the user does not speak directly, the robot can also detect the IoI if he or she looks at the robot for more than predefined periods of time. A state transition model for the proposed IoI detection framework is designed and verified by experiments with a mobile robot. In order to implement and associate our model in a robot architecture, all the components are implemented and integrated in the Robot Operating System(ROS) environment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。