用视觉语言模型捕捉非语言信号,让机器人更懂社交时机。
Using Vision-Language Models as Proxies for Social Intelligence in Human-Robot Interaction
- 用凝视转移和空间距离检测触发视觉语言模型推理
- 在真实场景回放中实现90%以上社交响应准确率
- 适合需要自然互动的家用或服务机器人研究
在大学咖啡厅开展为期五天的真人操控机器人部署实验,分析人们通过非语言行为表达互动意愿的方式,以及专家操控者如何据此判断交互时机。基于此,提出两阶段流水线:先用轻量级感知检测器(凝视转移、近身距离)识别社交关键时刻,再触发重型视频型视觉语言模型(VLM)进行深层语义理解。在回放的真实交互数据上评估该方法,对比两种提示策略,结果表明,仅在社会意义显著时刻调用VLM作为社交推理代理,可有效提升机器人响应的恰当性,使其能根据人们自然提供的线索做出合适行为。
原文摘要 · Abstract (English)
Robots operating in everyday environments must often decide when and whether to engage with people, yet such decisions often hinge on subtle nonverbal cues that unfold over time and are difficult to model explicitly. Drawing on a five-day Wizard-of-Oz deployment of a mobile service robot in a university cafe, we analyze how people signal interaction readiness through nonverbal behaviors and how expert wizards use these cues to guide engagement. Motivated by these observations, we propose a two-stage pipeline in which lightweight perceptual detectors (gaze shifts and proxemics) are used to selectively trigger heavier video-based vision-language model (VLM) queries at socially meaningful moments. We evaluate this pipeline on replayed field interactions and compare two prompting strategies. Our findings suggest that selectively using VLMs as proxies for social reasoning enables socially responsive robot behavior, allowing robots to act appropriately by attending to the cues people naturally provide in real-world interactions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。