arXiv:2410.01966cs.CVcs.AI2024-10被引 1

用多视角视觉语言模型提升儿童屏幕时间监测精度

Enhancing Screen Time Identification in Children with a Multi-View Vision Language Model and Screen Time Tracker

  • 融合可穿戴设备的多视角图像与视觉语言模型动态识别屏幕暴露
  • 在真实活动数据集上优于传统视觉模型和目标检测方法
  • 适合关注儿童行为研究与自然场景监测的研究者

准确监测儿童屏幕暴露对研究屏幕使用相关现象(如儿童肥胖、体力活动和社交互动)至关重要。现有研究多依赖自述或笨重可穿戴传感器的手动测量,效率与准确性不足。本文提出一种新型传感信息框架,利用可穿戴设备采集的中心视角图像(屏幕时间追踪器,STT)与视觉语言模型(VLM)。特别地,设计了一种多视角VLM,通过分析多个视角的中心视角图像序列,实现屏幕暴露的动态解析。在儿童自由活动数据集上的验证表明,该方法显著优于普通视觉语言模型与目标检测模型。结果证实了该监测方法在自然情境下优化儿童屏幕暴露行为研究的潜力。

原文摘要 · Abstract (English)

Being able to accurately monitor the screen exposure of young children is important for research on phenomena linked to screen use such as childhood obesity, physical activity, and social interaction. Most existing studies rely upon self-report or manual measures from bulky wearable sensors, thus lacking efficiency and accuracy in capturing quantitative screen exposure data. In this work, we developed a novel sensor informatics framework that utilizes egocentric images from a wearable sensor, termed the screen time tracker (STT), and a vision language model (VLM). In particular, we devised a multi-view VLM that takes multiple views from egocentric image sequences and interprets screen exposure dynamically. We validated our approach by using a dataset of children's free-living activities, demonstrating significant improvement over existing methods in plain vision language models and object detection models. Results supported the promise of this monitoring approach, which could optimize behavioral research on screen exposure in children's naturalistic settings.

儿童监测视觉语言模型多视角识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。