仅用一张照片估算食物份量,突破3D信息丢失难题
MFP3D: Monocular Food Portion Estimation Leveraging 3D Point Clouds
- 从单张图像重建3D点云,融合2D与3D特征
- 在MetaFood3D上体积估计误差降低18.7%,能量预测更准确
- 适合移动健康、饮食追踪等真实场景应用
食物份量估算是监测健康与追踪饮食摄入的关键。基于图像的饮食评估通过计算机视觉分析进食场景图像,正逐步取代传统的24小时回忆法。然而,由于3D信息在投影到2D图像平面时丢失,准确估算营养成分仍具挑战性。现有方法部署困难,依赖特定参考物、高质量深度图或多视角图像或视频。本文提出MFP3D,一种仅需单张单目图像即可实现精准食物份量估计的新框架。MFP3D包含三个核心模块:(1) 3D重建模块,从2D图像生成食物的3D点云表示;(2) 特征提取模块,融合3D点云与2D RGB图像特征;(3) 份量回归模块,利用深度回归模型根据提取特征估计食物体积与能量含量。在MetaFood3D数据集上的实验表明,MFP3D相较现有方法显著提升估计准确性。
原文摘要 · Abstract (English)
Food portion estimation is crucial for monitoring health and tracking dietary intake. Image-based dietary assessment, which involves analyzing eating occasion images using computer vision techniques, is increasingly replacing traditional methods such as 24-hour recalls. However, accurately estimating the nutritional content from images remains challenging due to the loss of 3D information when projecting to the 2D image plane. Existing portion estimation methods are challenging to deploy in real-world scenarios due to their reliance on specific requirements, such as physical reference objects, high-quality depth information, or multi-view images and videos. In this paper, we introduce MFP3D, a new framework for accurate food portion estimation using only a single monocular image. Specifically, MFP3D consists of three key modules: (1) a 3D Reconstruction Module that generates a 3D point cloud representation of the food from the 2D image, (2) a Feature Extraction Module that extracts and concatenates features from both the 3D point cloud and the 2D RGB image, and (3) a Portion Regression Module that employs a deep regression model to estimate the food's volume and energy content based on the extracted features. Our MFP3D is evaluated on MetaFood3D dataset, demonstrating its significant improvement in accurate portion estimation over existing methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。