用视觉语言模型自动标注可穿戴相机图像,减轻体力活动研究的人工标注负担。
Reducing Annotation Burden in Physical Activity Research Using Vision-Language Models
- 用开源视觉语言模型替代人工标注穿戴相机图像
- 对久坐行为预测准确率中位数达0.89,轻度及高强度活动下降至0.6左右
- 在跨地区数据上仍有效,适合需大规模标注的研究者
在自由生活环境中,可穿戴设备采集的数据若配有与健康研究兼容的体力活动标签,对验证现有测量方法和开发新机器学习模型至关重要。传统方式依赖人工逐帧标注参与者全天佩戴的摄像头拍摄图像,耗时费力。本研究对比了三种视觉语言模型(VLM)和两种判别模型(DM)在牛津郡(英国)和四川(中国)两个自由生活验证研究中的表现,数据来自使用Autographer可穿戴相机的161名和111名参与者。结果显示,最佳开源VLM与微调后的判别模型在牛津郡研究中对久坐行为预测性能相当,中位F1分数分别为0.89(0.84, 0.92)和0.91(0.86, 0.95);轻度活动分别为0.60(0.56, 0.67)和0.70(0.63, 0.79);中等至剧烈强度活动分别为0.66(0.53, 0.85)和0.72(0.58, 0.84)。在外部四川数据集上,性能普遍下降,中位Cohen's kappa分别从0.54(0.49, 0.64)降至0.26(0.15, 0.37),从0.67(0.60, 0.74)降至0.19(0.10, 0.30)。结论:免费可用的计算机视觉模型可用于类似人群的久坐行为标注,显著降低人工标注负担。
原文摘要 · Abstract (English)
Introduction: Data from wearable devices collected in free-living settings, and labelled with physical activity behaviours compatible with health research, are essential for both validating existing wearable-based measurement approaches and developing novel machine learning approaches. One common way of obtaining these labels relies on laborious annotation of sequences of images captured by cameras worn by participants through the course of a day. Methods: We compare the performance of three vision language models and two discriminative models on two free-living validation studies with 161 and 111 participants, collected in Oxfordshire, United Kingdom and Sichuan, China, respectively, using the Autographer (OMG Life, defunct) wearable camera. Results: We found that the best open-source vision-language model (VLM) and fine-tuned discriminative model (DM) achieved comparable performance when predicting sedentary behaviour from single images on unseen participants in the Oxfordshire study; median F1-scores: VLM = 0.89 (0.84, 0.92), DM = 0.91 (0.86, 0.95). Performance declined for light (VLM = 0.60 (0.56,0.67), DM = 0.70 (0.63, 0.79)), and moderate-to-vigorous intensity physical activity (VLM = 0.66 (0.53, 0.85); DM = 0.72 (0.58, 0.84)). When applied to the external Sichuan study, performance fell across all intensity categories, with median Cohen's kappa-scores falling from 0.54 (0.49, 0.64) to 0.26 (0.15, 0.37) for the VLM, and from 0.67 (0.60, 0.74) to 0.19 (0.10, 0.30) for the DM. Conclusion: Freely available computer vision models could help annotate sedentary behaviour, typically the most prevalent activity of daily living, from wearable camera images within similar populations to seen data, reducing the annotation burden.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。