arXiv:2607.05264cs.CV2026-07被引 1

评测视觉语言模型在真实工厂监控场景下的表现,发现现有模型严重不足。

SteelBench: Evaluating Vision-Language Models in Real-World Industrial Environments

论文配图:SteelBench: Evaluating Vision-Language Models in Real-World Industrial Environments
图 1 · 摘自论文原文
  • 基于149小时真实工厂视频构建1345段密集标注数据集
  • 最佳模型动作识别准确率仅42.6%,人类达84.6%
  • 首次引入标签溯源审计机制,揭示模型自标注的偏差风险

现有视频基准测试主要针对消费级视频、第一人称视角或模拟工业环境,未涵盖真实工厂闭路电视(CCTV)中的视觉与流程条件——如工人远距离出现、尘雾、低光、反光、遮挡及多重活动重叠。本文提出STEELBENCH,一个用于工业监控诊断的基准,联合评估个体作业行为识别、安全规则推理与标注溯源。该数据集包含1,345段密集标注视频片段,源自149小时实际厂区视频,通过时间去重、类别平衡与可见性感知分层采样,从10,024个候选片段中筛选得出。每段视频包含密集的逐人动作标签、个人防护装备(PPE)属性、空间上下文信息与安全规则标注。由于模型辅助标注可能影响后续评估所用标签,本研究引入溯源审计协议,测量标签影响力、评估对真实标签来源的敏感性,并提供专家审核的人类基准。经审计发现,未经审计的模型生成真值可使同类模型准确率虚高最高达17个百分点。在九种来自四种架构家族的视觉语言模型中,最佳模型动作识别准确率为42.6%,远低于人类基准84.6%。性能在识别、鲁棒性、校准与安全推理方面高度碎片化:即使预测正确动作,仍有37%-58%案例产生错误安全判断,且无一模型通过超过2项诊断检查。数据集已公开于Hugging Face。

原文摘要 · Abstract (English)

Existing video benchmarks evaluate action recognition on consumer videos, egocentric recordings, or simulated industrial environments. They do not test vision-language models under the visual and procedural conditions of real industrial CCTV, where workers appear as distant figures amid dust, steam, low light, glare, occlusion, and overlapping activities. We introduce STEELBENCH, a diagnostic benchmark for industrial surveillance that jointly evaluates per-worker activity recognition, safety-rule reasoning, and annotation provenance. SteelBench contains 1,345 densely annotated clips, curated from 149 hours of operational plant footage and 10,024 candidate clips using temporal deduplication, class balancing, and visibility-aware stratified sampling. Each clip includes dense per-worker action labels, PPE attributes, spatial context, and safety-rule annotations. Because model-assisted annotation can shape the labels later used for model evaluation, SteelBench includes a provenance-aware audit protocol. The protocol measures label influence, evaluates sensitivity to ground-truth provenance, and reports a human reference from expert-reviewed labels. Applying this audit, we find that unaudited VLM-sourced ground truth can inflate same-family model accuracy by up to 17 percentage points. Across nine VLMs from four architectural families, the best model reaches only 42.6% action accuracy, compared with an 84.6% human benchmark. Performance also fragments across recognition, robustness, calibration, and safety reasoning. Even when models predict the correct action, 37-58% of cases still yield incorrect safety judgments, and no model passes more than 2 of 5 diagnostic checks. The dataset is publicly available on Hugging Face.

工业视觉多模态评估安全检测数据溯源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。