用普通摄像头实现无需校准的自然交互,让机器人通过注视判断任务完成。
Gaze-Aware Task Progression Detection Framework for Human-Robot Interaction Using RGB Cameras
- 仅用单目摄像头和视觉算法追踪注视区域变化,识别任务进度。
- 任务完成检测准确率达77.6%,响应延迟略高但体验更自然舒适。
- 适合希望提升人机交互自然性的机器人研发人员或场景应用。
在人-机器人交互中,检测人类注视可帮助机器人理解用户注意力与意图。然而,多数注视检测方法依赖专用眼动追踪设备,限制了其在日常环境中的部署。基于外观的注视估计方法通过使用标准RGB摄像头消除了这一依赖,但在人-机器人交互中的实用性仍待探索。本文提出一种无需校准的框架,用于检测信息通过集成显示界面传递时的任务进展。该框架仅使用机器人内置单目RGB摄像头(640x480分辨率)和先进的注视估计技术,监测注意力模式。利用用户从任务界面转向机器人面部以示意任务完成的自然行为,定义三个兴趣区域(AOI):平板、机器人面部及其他区域。通过系统性参数优化,找到在检测精度与交互延迟间平衡的最佳配置。我们在“第一天上班”场景中验证该框架,并与按键交互进行对比。结果表明,任务完成检测准确率为77.6%。相较于按键交互,本系统响应延迟稍高,但保留了信息记忆能力,显著提升了舒适度、社会存在感及感知自然性。值得注意的是,多数参与者表示并未有意识地用眼神引导交互,凸显注视作为沟通线索的直观性。本工作展示了低成本、纯RGB摄像头实现自然、沉浸式人机交互的可行性。
原文摘要 · Abstract (English)
In human-robot interaction (HRI), detecting a human's gaze helps robots interpret user attention and intent. However, most gaze detection approaches rely on specialized eye-tracking hardware, limiting deployment in everyday settings. Appearance-based gaze estimation methods remove this dependency by using standard RGB cameras, but their practicality in HRI remains underexplored. We present a calibration-free framework for detecting task progression when information is conveyed via integrated display interfaces. The framework uses only the robot's built-in monocular RGB camera (640x480 resolution) and state-of-the-art gaze estimation to monitor attention patterns. It leverages natural behavior, where users shift focus from task interfaces to the robot's face to signal task completion, formalized through three Areas of Interest (AOI): tablet, robot face, and elsewhere. Systematic parameter optimization identifies configurations that balance detection accuracy and interaction latency. We validate our framework in a "First Day at Work" scenario, comparing it to button-based interaction. Results show a task completion detection accuracy of 77.6%. Compared to button-based interaction, the proposed system exhibits slightly higher response latency but preserves information retention and significantly improves comfort, social presence, and perceived naturalness. Notably, most participants reported that they did not consciously use eye movements to guide the interaction, underscoring the intuitive role of gaze as a communicative cue. This work demonstrates the feasibility of intuitive, low-cost, RGB-only gaze-based HRI for natural and engaging interactions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。