用大脑看东西的原理,让纸板机器人实现不需传感器的对视。
Perception Is All You Need: A Neuroscience Framework for Low Cost Sensorless Gaze in HRI

- 利用人脑对凸面的固有感知,设计可低成本制作的纸板机器人。
- 实验表明凹面眼窝加画瞳孔即可在任意角度引发对视错觉。
- 适合教育、儿童交互等场景,无需电力与隐私保护措施。
在儿童-机器人互动中,注视跟随能提升注意力、记忆和学习效果,但现有方案需花费超过3万美元,依赖昂贵传感器、算法,并引发隐私问题。本文提出一种无需传感器与计算的框架,基于人类视觉系统对凸面的先验假设,实现机器人与观察者间的感知式注视跟随。具体而言,通过设计单价低于一美元的纸板机器人,反向模拟大脑的注视计算机制,使观察者的感知系统成为机器人的“执行器”。该框架建立在三方面神经科学证据之上:大脑通过上颞沟处理面部注视方向;高精度的凸面优先机制使人将凹面误认为凸面;预测加工层级中高层面部知识会覆盖底层深度信号。这些机制解释了为何凹面眼窝加画瞳孔可在任何视角产生相互凝视的错觉。我们从感知科学推导出设计约束,提供一款开源可替换眼罩的亚美元级机器人原型,并识别出成功与失败的边界条件(发展、临床与几何)。若被广泛应用,过去二十年的HRI注视研究发现有望实现大规模落地。
原文摘要 · Abstract (English)
Gaze-following in child-robot interaction improves attention, recall, and learning, but requires expensive platforms (\$30,000+), sensors, algorithms, and raises privacy concerns. We propose a framework that avoids sensors and computation entirely, instead relying on the human visual system's assumption of convexity to produce perceptual gaze-following between a robot and its viewer. Specifically, we motivate sub-dollar cardboard robot design that directly implements the brain's own gaze computation pipeline in reverse, making the viewer's perceptual system the robot's "actuator", with no sensors, no power, and no privacy concerns. We ground this framework in three converging lines of theoretical and empirical neuroscience evidence. Namely, the distributed face processing network that computes gaze direction via the superior temporal sulcus, the high-precision convexity prior that causes the brain to perceive concave faces as convex, and the predictive processing hierarchy in which top-down face knowledge overrides bottom-up depth signals. These mechanisms explain why a concave eye socket with a painted pupil produces the perception of mutual gaze from any viewing angle. We derive design constraints from perceptual science, present a sub-dollar open-template robot with parameterized interchangeable eye inserts, and identify boundary conditions (developmental, clinical, and geometric) that predict where the framework will succeed and where it will fail. If leveraged, two decades of HRI gaze findings become deliverable at population scale.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。