用视觉语言模型让机器人自适应跟拍动态变队的人群,更自然安全。
Adaptive Companionship for Group-Following Robots: Handling Dynamically Changing Group Formations

- 用VLM理解人群动态,推断成员位置与社交距离
- 实验显示成功率提升15%,碰撞率降低25%
- 适合开发具社会智能的陪伴机器人
伴随人类群体是机器人发展类人社交认知的关键。然而,人类群体通常不保持固定队形,给机器人维持自然陪伴行为带来挑战。本文提出一种基于视觉-语言模型(VLMs)的自适应群体伴行方法,利用其语义推理能力推断成员位置、保持社交距离并理解群体动态。首先检测群体成员,感知模块生成交互空间的视觉表征作为VLM输入,并与模型预测路径积分(MPPI)控制器结合,确保系统稳定与安全。在五个场景下的实验表明,该方法显著提升伴行效果,相较于基线方法成功率达15%提升,碰撞率降低25%。此外,用户研究显示生成的陪伴行为被感知为自然且符合社交规范。
原文摘要 · Abstract (English)
Accompanying a group of humans is an essential aspect of developing human-like social cognition in robots. However, human groups typically do not follow fixed formations, which poses significant challenges for robots in maintaining natural companionship behaviors. In this paper, we propose an adaptive group-accompaniment method for social robots based on Vision-Language Models (VLMs), leveraging their semantic reasoning capabilities to infer companion positions, maintain social distances, and understand group dynamics. The members of the group are first detected, and a perceptual module generates visual representations of the interaction group space as input to the VLM, which is then combined with a Model Predictive Path Integral (MPPI) controller to ensure stability and safety. Experimental evaluations across five scenarios show that the proposed method enables robots to accompany the group effectively, demonstrating a 15\% improvement in success rate and a 25\% reduction in collision rate compared to baseline approaches. Additionally, a user study indicates that the generated companionship behaviors are perceived as natural and socially appropriate.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。