用视觉语言模型从视频中自动估算搬运任务的双手距离,助力工伤风险评估。
Vision-Language Models for Ergonomic Assessment of Manual Lifting Tasks: Estimating Horizontal and Vertical Hand Distances from RGB Video
- 通过文本引导的视觉定位与时序回归,从视频中非侵入式估计手部水平与垂直距离。
- 多视角分割管道误差最小,水平距离误差约6-8厘米,垂直距离误差5-8厘米。
- 适合工业安全、人因工程领域,尤其在无传感器环境中评估搬运风险。
手动搬运是导致职业性肌肉骨骼疾病的主要因素,有效的人体工学风险评估对于量化身体暴露并指导干预措施至关重要。修订版NIOSH搬运方程(RNLE)是广泛使用的人体工学风险评估工具,依赖六个任务变量,包括水平(H)和垂直(V)手部距离;这些距离通常需人工测量或借助专用传感系统获取,在真实环境中的应用存在困难。本文评估了创新性视觉语言模型(VLMs)从RGB视频流中非侵入式估算H和V的可行性。开发了两种多阶段VLM管道:仅检测的文本引导管道与检测加分割管道。两者均采用文本引导定位任务相关区域,提取视觉特征,并通过Transformer进行时序回归,以估算搬运起始与结束时的H和V。在多种搬运任务下,采用留一被试者交叉验证,评估两个管道及七种相机视角条件下的性能。结果在不同管道和视角间差异显著,基于分割的多视角管道始终表现最佳,估算水平距离的平均绝对误差约为6-8厘米,垂直距离为5-8厘米。相较仅检测管道,像素级分割使水平距离误差降低约20-30%,垂直距离降低35-40%。研究结果支持基于VLM的管道用于视频驱动的RNLE参数估算。
原文摘要 · Abstract (English)
Manual lifting tasks are a major contributor to work-related musculoskeletal disorders, and effective ergonomic risk assessment is essential for quantifying physical exposure and informing ergonomic interventions. The Revised NIOSH Lifting Equation (RNLE) is a widely used ergonomic risk assessment tool for lifting tasks that relies on six task variables, including horizontal (H) and vertical (V) hand distances; such distances are typically obtained through manual measurement or specialized sensing systems and are difficult to use in real-world environments. We evaluated the feasibility of using innovative vision-language models (VLMs) to non-invasively estimate H and V from RGB video streams. Two multi-stage VLM-based pipelines were developed: a text-guided detection-only pipeline and a detection-plus-segmentation pipeline. Both pipelines used text-guided localization of task-relevant regions of interest, visual feature extraction from those regions, and transformer-based temporal regression to estimate H and V at the start and end of a lift. For a range of lifting tasks, estimation performance was evaluated using leave-one-subject-out validation across the two pipelines and seven camera view conditions. Results varied significantly across pipelines and camera view conditions, with the segmentation-based, multi-view pipeline consistently yielding the smallest errors, achieving mean absolute errors of approximately 6-8 cm when estimating H and 5-8 cm when estimating V. Across pipelines and camera view configurations, pixel-level segmentation reduced estimation error by approximately 20-30% for H and 35-40% for V relative to the detection-only pipeline. These findings support the feasibility of VLM-based pipelines for video-based estimation of RNLE distance parameters.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。