arXiv:2608.22586cs.CVcs.AI2026-08

用RGB视频和物体重量估算搬运时双手受力,无需穿戴设备。

Vision-Language Models for Occupational Physical Exposure Assessment: Estimating External Hand Forces in Manual Material Handling Tasks from RGB Video

  • 结合文本提示与视觉特征,从视频中定位人和物体区域,预测三轴受力。
  • 水平/内外向力误差约4.7-5.6牛,垂直力误差约10.6-11.0牛。
  • 单摄像头下加物体区域可提升精度,多视角对峰值力估计最有益。

外部手部受力是职业体力暴露和损伤风险生物力学分析的重要输入,但传统连续测量需使用传感工具或特殊传感器。本文评估了一种基于视觉语言模型(VLM)的流程,结合任务特定文本提示、视觉表征及已知箱体质量,从RGB视频中估计动态、三轴、双侧外部手部受力。35名健康年轻成人完成5种搬运任务(提、搬、推、拉),使用6、9、12公斤箱体。该流程采用文本引导的区域定位、预训练视觉变换器特征提取及基于Transformer的时间序列回归。在七种相机视角条件(三种单视角、四种多视角)和四种感兴趣区域(ROI)策略下进行留一被试者验证。总体均方根误差为:水平/内外向力约4.7-5.6牛,垂直力约10.6-11.0牛。将被搬运物体作为第二个ROI通常提升估计效果,尤其在单相机条件下;而像素级分割带来的改进有限。多相机采集在峰值力估计上优势明显,尤其是垂直分量;不同相机配置的整体帧级误差差异较小。结果表明,仅凭RGB视频与已知载荷质量即可实现连续、双侧、方向性手力估计,无需在人员或物体上安装传感器,支持更可扩展的职业体力暴露与风险评估发展。

原文摘要 · Abstract (English)

External hand forces are important inputs to biomechanical analyses of occupational physical exposure and injury risk, yet continuous force measurements during manual material handling (MMH) typically requires instrumented objects or specialized sensing. We evaluated a vision-language model (VLM)-based pipeline that combines task-specific textual cues, visual representations, and known box mass to estimate dynamic, triaxial, bilateral external hand forces from RGB video. Thirty-five healthy young adults performed five MMH tasks involving lifting, carrying, pushing, and pulling with box masses of 6, 9, and 12 kg. The pipeline used text-guided localization of participant and handled-object regions of interest (ROIs), pretrained vision-transformer feature extraction, and transformer-based temporal regression. Performance was evaluated using leave-one-subject-out validation across seven camera-view conditions (three single-view and four multi-view conditions) and four ROI strategies. Overall, root mean square error was ~4.7-5.6 N for the horizontal and mediolateral force components and ~10.6-11.0 N for the vertical component. Including the handled object as a second ROI generally improved force estimation, with some of the largest benefits under single-camera conditions, whereas pixel-level segmentation provided little additional improvement. Multi-camera capture provided the clearest benefit for peak-force estimation, particularly for the vertical component, whereas differences in overall frame-level error among camera configurations were comparatively modest. These findings demonstrate the feasibility of estimating continuous, bilateral, directional hand-force estimates from RGB video and known load mass without requiring sensors on the worker or handled objects as model inputs, supporting the development of more scalable occupational physical exposure and risk assessments.

手部受力估计视觉语言模型搬运任务无传感器测量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。