arXiv:2602.10946cs.RO2026-02

用深度学习让机器人学会自然看人,提升社交互动真实感。

Developing Neural Network-Based Gaze Control Systems for Social Robots

  • 用LSTM和Transformer分析眼动数据,建模人类在社交场景中的注视模式。
  • 2D动画预测准确率60%,3D动画达65%,落地到Nao机器人后获用户好评。
  • 适合想提升人机交互自然性的机器人开发者参考。

在多人互动中,视线方向是兴趣与意图的关键信号,因此社交机器人需能恰当地转移注意力。理解社交情境对机器人有效参与、预测人类意图及顺畅导航互动至关重要。本研究基于30名参与者的数据,利用深度神经网络构建了人类在多种社交情境(如进入、离开、挥手、说话、指向)下的眼动行为经验模式。我们制作了两段视频:一段用于电脑屏幕,另一段用于虚拟现实头显,呈现不同社交场景。数据采集自15名使用眼动仪的参与者和15名使用Oculus Quest 1头显的参与者。采用长短期记忆网络(LSTM)和Transformer模型分析并预测注视模式。模型在2D动画中实现60%的注视方向预测准确率,在3D动画中达65%。随后将最优模型部署至Nao机器人,36名新参与者进行评估。反馈显示整体满意度较高,有机器人经验者评价更积极。

原文摘要 · Abstract (English)

During multi-party interactions, gaze direction is a key indicator of interest and intent, making it essential for social robots to direct their attention appropriately. Understanding the social context is crucial for robots to engage effectively, predict human intentions, and navigate interactions smoothly. This study aims to develop an empirical motion-time pattern for human gaze behavior in various social situations (e.g., entering, leaving, waving, talking, and pointing) using deep neural networks based on participants' data. We created two video clips-one for a computer screen and another for a virtual reality headset-depicting different social scenarios. Data were collected from 30 participants: 15 using an eye-tracker and 15 using an Oculus Quest 1 headset. Deep learning models, specifically Long Short-Term Memory (LSTM) and Transformers, were used to analyze and predict gaze patterns. Our models achieved 60% accuracy in predicting gaze direction in a 2D animation and 65% accuracy in a 3D animation. Then, the best model was implemented onto the Nao robot; and 36 new participants evaluated its performance. The feedback indicated overall satisfaction, with those experienced in robotics rating the models more favorably.

社交机器人眼动预测LSTMTransformer

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。