让机器人学会像人一样看东西,包括对非人类动作的反应。
Human-Like Gaze Behavior in Social Robots: A Deep Learning Approach Integrating Human and Non-Human Stimuli
- 用LSTM和Transformer模型预测人类在真实与虚拟场景中的注视方向。
- 在真实场景中模型准确率达72%,优于以往方法。
- 首次考虑非人类刺激(如开门、掉物)对眼神的影响,适合社交机器人研发者。
非语言行为,尤其是视线方向,在提升社交互动有效性中起关键作用。随着社交机器人参与越来越多的互动,它们必须根据人类活动调整视线,并对所有线索(无论是否由人类产生)保持敏感,以确保沟通顺畅有效。本研究旨在提升机器人与人在各种社交情境下注视行为的相似性,涵盖人类与非人类刺激(如对话、指向、门开启、物体掉落)。一个关键创新在于探讨对非人类刺激的注视响应,这是此前研究较少关注的领域。这些情景在Unity软件中通过3D动画和360度现实视频进行模拟,利用虚拟现实眼镜收集了41名参与者的眼动数据。预处理后,训练了LSTM和Transformer两种神经网络构建预测模型。在动画场景中,LSTM与Transformer模型的预测准确率分别为67.6%和70.4%;在真实世界场景中,准确率分别为72%和71.6%。尽管个体间注视模式存在差异,但本模型在准确性上优于现有方法,且独特地纳入了非人类刺激。此外,在NAO机器人上部署该系统,经275名参与者问卷评估,交互满意度高。本工作推动了社交机器人技术发展,使机器人能在复杂社交环境中动态模仿人类的注视行为。
原文摘要 · Abstract (English)
Nonverbal behaviors, particularly gaze direction, play a crucial role in enhancing effective communication in social interactions. As social robots increasingly participate in these interactions, they must adapt their gaze based on human activities and remain receptive to all cues, whether human-generated or not, to ensure seamless and effective communication. This study aims to increase the similarity between robot and human gaze behavior across various social situations, including both human and non-human stimuli (e.g., conversations, pointing, door openings, and object drops). A key innovation in this study, is the investigation of gaze responses to non-human stimuli, a critical yet underexplored area in prior research. These scenarios, were simulated in the Unity software as a 3D animation and a 360-degree real-world video. Data on gaze directions from 41 participants were collected via virtual reality (VR) glasses. Preprocessed data, trained two neural networks-LSTM and Transformer-to build predictive models based on individuals' gaze patterns. In the animated scenario, the LSTM and Transformer models achieved prediction accuracies of 67.6% and 70.4%, respectively; In the real-world scenario, the LSTM and Transformer models achieved accuracies of 72% and 71.6%, respectively. Despite the gaze pattern differences among individuals, our models outperform existing approaches in accuracy while uniquely considering non-human stimuli, offering a significant advantage over previous literature. Furthermore, deployed on the NAO robot, the system was evaluated by 275 participants via a comprehensive questionnaire, with results demonstrating high satisfaction during interactions. This work advances social robotics by enabling robots to dynamically mimic human gaze behavior in complex social contexts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。