构建工业安全视频语言评测基准,检验模型在真实场景下的安全识别能力
iSafetyBench: A video-language benchmark for safety in industrial environment
- 基于1100段真实工业视频,标注98类常规与67类危险动作
- 多模型零样本测试显示,识别危险动作准确率普遍低于40%
- 适合研究工业安全、多模态模型鲁棒性的团队使用
视觉语言模型(VLMs)在零样本视频理解任务中展现出强大泛化能力,但在高风险工业场景中——需同时识别常规操作与安全关键异常——其表现仍不明确。为此,我们提出iSafetyBench,首个专为工业环境设计的视频-语言评测基准,涵盖正常与危险场景。该基准包含1,100段真实工业视频,标注了98类常规动作和67类危险动作,采用开放词汇、多标签标签体系,并配有多选题用于单标签与多标签评估。我们在零样本条件下测试了8个先进VLM模型。尽管在现有基准上表现优异,这些模型在iSafetyBench上表现不佳,尤其在识别危险行为和多标签场景中准确率显著下降。结果揭示了巨大性能差距,凸显开发更可靠、安全感知的多模态模型的迫切需求。iSafetyBench为推动该领域发展提供了首个综合性测试平台。数据集可从https://github.com/iSafetyBench/data获取。
原文摘要 · Abstract (English)
Recent advances in vision-language models (VLMs) have enabled impressive generalization across diverse video understanding tasks under zero-shot settings. However, their capabilities in high-stakes industrial domains-where recognizing both routine operations and safety-critical anomalies is essential-remain largely underexplored. To address this gap, we introduce iSafetyBench, a new video-language benchmark specifically designed to evaluate model performance in industrial environments across both normal and hazardous scenarios. iSafetyBench comprises 1,100 video clips sourced from real-world industrial settings, annotated with open-vocabulary, multi-label action tags spanning 98 routine and 67 hazardous action categories. Each clip is paired with multiple-choice questions for both single-label and multi-label evaluation, enabling fine-grained assessment of VLMs in both standard and safety-critical contexts. We evaluate eight state-of-the-art video-language models under zero-shot conditions. Despite their strong performance on existing video benchmarks, these models struggle with iSafetyBench-particularly in recognizing hazardous activities and in multi-label scenarios. Our results reveal significant performance gaps, underscoring the need for more robust, safety-aware multimodal models for industrial applications. iSafetyBench provides a first-of-its-kind testbed to drive progress in this direction. The dataset is available at: https://github.com/iSafetyBench/data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。