用6D姿态估计追踪手和勺子,实现饮食行为的精准监测。
6D Pose Estimation on Spoons and Hands
- 基于视频流,通过6D姿态估计追踪手与勺子的空间位置和朝向。
- 对比两种先进视频分割模型,验证系统在真实场景下的性能表现。
- 为饮食行为分析提供可量化依据,适用于营养研究与健康干预。
精准的饮食监测对促进健康饮食习惯至关重要。研究重点在于人们使用餐具和手进食时的交互行为。通过跟踪其位置与姿态,可估算食物摄入量或监测进食行为,为营养摄入提供比自述法更可靠的洞察。本文构建了一个系统,利用6D姿态估计分析静止视频中进食者的动作,实时追踪手和勺子的运动。我们定量与定性地评估了两种前沿视频对象分割(VOS)模型的表现,并识别出系统中的主要误差来源。
原文摘要 · Abstract (English)
Accurate dietary monitoring is essential for promoting healthier eating habits. A key area of research is how people interact and consume food using utensils and hands. By tracking their position and orientation, it is possible to estimate the volume of food being consumed, or monitor eating behaviours, highly useful insights into nutritional intake that can be more reliable than popular methods such as self-reporting. Hence, this paper implements a system that analyzes stationary video feed of people eating, using 6D pose estimation to track hand and spoon movements to capture spatial position and orientation. In doing so, we examine the performance of two state-of-the-art (SOTA) video object segmentation (VOS) models, both quantitatively and qualitatively, and identify main sources of error within the system.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。