用下巴摄像头捕捉旁观者表情,让机器人识别错误并自我修正。
"Why the face?": Exploring Robot Error Detection Using Instrumented Bystander Reactions
- 通过下巴佩戴设备采集旁观者面部微表情,重建3D人脸动作。
- 新模型在个体数据上表现优于OpenFace等传统方法。
- 适合希望提升机器人社交感知能力的研究者与开发者。
人类如何察觉并纠正社交失误?我们通过观察同伴的微妙反应——如挑眉、轻笑——来评估环境与自身行为。然而,机器人难以感知并利用这些细微线索。本文提出一种新型颈挂式装置,从下巴区域记录面部表情,探索此前未被利用的数据以捕捉人类对机器人错误的反应。首先,我们开发了NeckNet-18,一个3D面部重建模型,将下巴摄像头捕捉到的反应映射为面部关键点与头部运动。随后,基于这些面部反应构建机器人错误检测模型,在跨参与者与同参与者数据上均优于OpenFace或视频数据基准,尤其在个体数据中表现更优。本研究倡导扩展人机协同的机器人感知,推动机器人在多样人类环境中更自然地融入,拓展社会线索检测边界,开辟适应性机器人新路径。
原文摘要 · Abstract (English)
How do humans recognize and rectify social missteps? We achieve social competence by looking around at our peers, decoding subtle cues from bystanders - a raised eyebrow, a laugh - to evaluate the environment and our actions. Robots, however, struggle to perceive and make use of these nuanced reactions. By employing a novel neck-mounted device that records facial expressions from the chin region, we explore the potential of previously untapped data to capture and interpret human responses to robot error. First, we develop NeckNet-18, a 3D facial reconstruction model to map the reactions captured through the chin camera onto facial points and head motion. We then use these facial responses to develop a robot error detection model which outperforms standard methodologies such as using OpenFace or video data, generalizing well especially for within-participant data. Through this work, we argue for expanding human-in-the-loop robot sensing, fostering more seamless integration of robots into diverse human environments, pushing the boundaries of social cue detection and opening new avenues for adaptable robotics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。