用人体骨骼数据识别吃喝动作,准确率超85%
Skeleton-Based Intake Gesture Detection With Spatial-Temporal Graph Convolutional Networks
- 结合空时图卷积与双向LSTM,从骨骼序列中捕捉吃喝动作特征
- 在实验室和手机拍摄场景下,吃喝动作识别F1值分别达86.18%和85.40%
- 适合智能健康监测、隐私敏感场景下的饮食行为分析
肥胖与超重已成为普遍社会问题,常与不健康饮食模式相关。通过自动识别进食动作来提升日常饮食监测是一种有前景的方法。本文提出一种基于骨骼信息的检测方法,融合空时图卷积网络(ST-GCN)与双向长短期记忆网络(BiLSTM),构建ST-GCN-BiLSTM模型,用于识别摄入动作。该方法具备环境鲁棒性强、数据依赖低、隐私保护优等优势。采用两个数据集验证模型性能:在实验室录制的OREBA数据集上,吃和喝动作的片段级F1得分分别为86.18%和74.84%;在自收集的手机拍摄数据集上,使用OREBA训练的模型分别取得85.40%和67.80%的F1得分。结果表明,利用骨骼数据进行摄入动作检测具有可行性,且模型在跨数据集验证中表现出良好鲁棒性。
原文摘要 · Abstract (English)
Overweight and obesity have emerged as widespread societal challenges, frequently linked to unhealthy eating patterns. A promising approach to enhance dietary monitoring in everyday life involves automated detection of food intake gestures. This study introduces a skeleton based approach using a model that combines a dilated spatial-temporal graph convolutional network (ST-GCN) with a bidirectional long-short-term memory (BiLSTM) framework, as called ST-GCN-BiLSTM, to detect intake gestures. The skeleton-based method provides key benefits, including environmental robustness, reduced data dependency, and enhanced privacy preservation. Two datasets were employed for model validation. The OREBA dataset, which consists of laboratory-recorded videos, achieved segmental F1-scores of 86.18% and 74.84% for identifying eating and drinking gestures. Additionally, a self-collected dataset using smartphone recordings in more adaptable experimental conditions was evaluated with the model trained on OREBA, yielding F1-scores of 85.40% and 67.80% for detecting eating and drinking gestures. The results not only confirm the feasibility of utilizing skeleton data for intake gesture detection but also highlight the robustness of the proposed approach in cross-dataset validation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。