arXiv:2504.10873cs.CVcs.AI2025-04CVPR被引 10

测试大模型能否零样本理解行人交通手势,发现现有模型表现差,难用于自动驾驶。

Can Vision-Language Models Understand and Interpret Dynamic Gestures from Pedestrians? Pilot Datasets and Exploration Towards Instructive Nonverbal Commands for Cooperative Autonomous Vehicles

  • 构建两个行人手势数据集,含正式与非正式指令,配自然语言标注。
  • 模型在手势分类上F1仅0.14-0.39,句子相似度低于0.59,远低于专家水平。
  • 虽姿态重建有潜力,但需更多数据和更好评估方法,适合自动驾驶研究者。

在自动驾驶中,正确解读交通手势(TGs)——如权威人员指令或行人示意——对保障所有道路使用者的安全与舒适至关重要。本研究探究了当前最先进的视觉语言模型(VLMs)在零样本情境下对交通场景中人类手势的描述与分类能力。我们构建并公开发布两个定制数据集:'表演式交通手势(ATG)'与'野外指导性交通手势(ITGI)',涵盖'停止'、'倒车'、'招手'等不同正式与非正式手势,并以自然语言标注行人体位与动作。通过三种方法评估模型:(1)生成句与专家句的相似度,(2)手势分类准确率,(3)姿态序列重建相似度。结果表明,当前VLMs在手势理解方面表现不佳:句子相似度平均低于0.59,分类F1得分仅为0.14–0.39,显著低于专家基线的0.70。尽管姿态重建显示出一定潜力,但需更多数据与更精细的评估指标才能可靠应用。研究揭示,尽管部分前沿VLM可进行零样本手势理解,但尚无模型具备足够的准确性与鲁棒性,亟需该领域深入研究。

原文摘要 · Abstract (English)

In autonomous driving, it is crucial to correctly interpret traffic gestures (TGs), such as those of an authority figure providing orders or instructions, or a pedestrian signaling the driver, to ensure a safe and pleasant traffic environment for all road users. This study investigates the capabilities of state-of-the-art vision-language models (VLMs) in zero-shot interpretation, focusing on their ability to caption and classify human gestures in traffic contexts. We create and publicly share two custom datasets with varying formal and informal TGs, such as 'Stop', 'Reverse', 'Hail', etc. The datasets are "Acted TG (ATG)" and "Instructive TG In-The-Wild (ITGI)". They are annotated with natural language, describing the pedestrian's body position and gesture. We evaluate models using three methods utilizing expert-generated captions as baseline and control: (1) caption similarity, (2) gesture classification, and (3) pose sequence reconstruction similarity. Results show that current VLMs struggle with gesture understanding: sentence similarity averages below 0.59, and classification F1 scores reach only 0.14-0.39, well below the expert baseline of 0.70. While pose reconstruction shows potential, it requires more data and refined metrics to be reliable. Our findings reveal that although some SOTA VLMs can interpret zero-shot human traffic gestures, none are accurate and robust enough to be trustworthy, emphasizing the need for further research in this domain.

视觉语言模型自动驾驶手势识别零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。