用第一视角视频直接测量手持食物体积,不依赖手势和固定咬量假设。
FoodTrack: Estimating Handheld Food Portions with Egocentric Video
- 基于第一视角视频,直接估算手持食物体积
- 绝对误差约7.01%,优于之前16.40%的最好结果
- 对手部遮挡和相机角度变化具有鲁棒性
准确追踪食物摄入对营养与健康监测至关重要。传统方法通常依赖特定拍摄角度、无遮挡图像或手势识别来估算摄入量,往往基于固定咬量假设,而非直接测量食物体积。本文提出FoodTrack框架,利用第一视角视频追踪并测量手持食物体积,对手部遮挡和相机/物体姿态变化具有鲁棒性。该方法直接估计食物体积,无需依赖进食手势或固定咬量假设,提供更准确、更灵活的解决方案。在手持食物对象上,绝对百分比误差约为7.01%,优于此前在更受限条件下达到16.40%平均绝对百分比误差的方法。
原文摘要 · Abstract (English)
Accurately tracking food consumption is crucial for nutrition and health monitoring. Traditional approaches typically require specific camera angles, non-occluded images, or rely on gesture recognition to estimate intake, making assumptions about bite size rather than directly measuring food volume. We propose the FoodTrack framework for tracking and measuring the volume of hand-held food items using egocentric video which is robust to hand occlusions and flexible with varying camera and object poses. FoodTrack estimates food volume directly, without relying on intake gestures or fixed assumptions about bite size, offering a more accurate and adaptable solution for tracking food consumption. We achieve absolute percentage loss of approximately 7.01% on a handheld food object, improving upon a previous approach that achieved a 16.40% mean absolute percentage error in its best case, under less flexible conditions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。