让机器人像人一样自然转头看物,提升人机互动真实感。
Humanizing Robot Gaze Shifts: A Framework for Natural Gaze Shifts in Humanoid Robots
- 用视觉语言模型分析多模态线索,判断该往哪看。
- 通过条件量化变分自编码器生成多样且自然的视线移动动作。
- 适合研究人机交互、具身智能与仿生运动控制的学者。
在社交互动中,利用听觉和视觉反馈实现注意力重新定向对自然的眼神转移至关重要。然而,在非约束性人机交互(HRI)中,使类人机器人实现自然且符合情境的眼神转移仍面临挑战,这需要认知注意力机制与生物仿生运动生成的协同。本文提出机器人眼神转移(RGS)框架,将这两个组件整合为统一流程。首先,RGS采用基于视觉-语言模型(VLM)的眼部推理管道,从多模态交互线索中推断出符合情境的眼神目标,确保与人类眼神指向规律一致。其次,RGS引入一种条件向量量化变分自编码器(VQ-VAE)模型,用于生成眼-头协同的眼神转移动作,产生多样化且类人的行为。实验验证了RGS能有效复现人类般的目标选择,并生成真实、多样的眼神转移动作。
原文摘要 · Abstract (English)
Leveraging auditory and visual feedback for attention reorientation is essential for natural gaze shifts in social interaction. However, enabling humanoid robots to perform natural and context-appropriate gaze shifts in unconstrained human--robot interaction (HRI) remains challenging, as it requires the coupling of cognitive attention mechanisms and biomimetic motion generation. In this work, we propose the Robot Gaze-Shift (RGS) framework, which integrates these two components into a unified pipeline. First, RGS employs a vision--language model (VLM)-based gaze reasoning pipeline to infer context-appropriate gaze targets from multimodal interaction cues, ensuring consistency with human gaze-orienting regularities. Second, RGS introduces a conditional Vector Quantized-Variational Autoencoder (VQ-VAE) model for eye--head coordinated gaze-shift motion generation, producing diverse and human-like gaze-shift behaviors. Experiments validate that RGS effectively replicates human-like target selection and generates realistic, diverse gaze-shift motions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。