arXiv:2601.10835cs.CVcs.AI2026-01被引 2

测试三大视觉语言模型在工地行为与情绪识别中的表现,发现GPT-4o最出色。

Can Vision-Language Models Understand Construction Workers? An Exploratory Study

  • 用三款主流VLM模型分析1000张工地图像中的动作与情绪。
  • GPT-4o在动作识别上F1达0.756,情绪识别F1为0.712,领先其他模型。
  • 模型难区分相似动作,适合做初步评估,需改进才可用于实际场景。

随着机器人在建筑流程中日益普及,其理解并响应人类行为的能力对实现安全高效协作至关重要。视觉语言模型(VLMs)在视觉理解任务中展现出潜力,可在无需大量领域特定训练的情况下识别人类行为,尤其适用于标签数据稀缺的建筑领域。本研究评估了GPT-4o、Florence 2和LLaVa-1.5三款领先VLM在静态工地图像中检测工人动作与情绪的表现。基于包含1000张图像、标注了十类动作与十类情绪的自建数据集,采用标准化推理流程与多评估指标进行测试。GPT-4o在两项任务中均表现最优,动作识别平均F1得分为0.756,准确率为0.799;情绪识别F1为0.712,准确率为0.773。Florence 2表现中等,动作与情绪识别的F1分别为0.497和0.414;LLaVa-1.5表现最差,两项任务的F1分别为0.466和0.461。混淆矩阵分析显示,所有模型均难以区分语义相近类别,如团队协作与向主管沟通。结果表明,通用VLM可为建筑环境提供人类行为识别的基线能力,但要实现真实可靠性,仍需领域适配、时序建模或多模态感知等改进。

原文摘要 · Abstract (English)

As robotics become increasingly integrated into construction workflows, their ability to interpret and respond to human behavior will be essential for enabling safe and effective collaboration. Vision-Language Models (VLMs) have emerged as a promising tool for visual understanding tasks and offer the potential to recognize human behaviors without extensive domain-specific training. This capability makes them particularly appealing in the construction domain, where labeled data is scarce and monitoring worker actions and emotional states is critical for safety and productivity. In this study, we evaluate the performance of three leading VLMs, GPT-4o, Florence 2, and LLaVa-1.5, in detecting construction worker actions and emotions from static site images. Using a curated dataset of 1,000 images annotated across ten action and ten emotion categories, we assess each model's outputs through standardized inference pipelines and multiple evaluation metrics. GPT-4o consistently achieved the highest scores across both tasks, with an average F1-score of 0.756 and accuracy of 0.799 in action recognition, and an F1-score of 0.712 and accuracy of 0.773 in emotion recognition. Florence 2 performed moderately, with F1-scores of 0.497 for action and 0.414 for emotion, while LLaVa-1.5 showed the lowest overall performance, with F1-scores of 0.466 for action and 0.461 for emotion. Confusion matrix analyses revealed that all models struggled to distinguish semantically close categories, such as collaborating in teams versus communicating with supervisors. While the results indicate that general-purpose VLMs can offer a baseline capability for human behavior recognition in construction environments, further improvements, such as domain adaptation, temporal modeling, or multimodal sensing, may be needed for real-world reliability.

视觉语言模型工地安全行为识别AI评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。