从单张图片重建真实尺寸的3D食物模型,提升饮食量估算精度。
Size Matters: Reconstructing Real-Scale 3D Models from Monocular Images for Food Portion Estimation
- 利用大规模预训练模型提取视觉特征,估计物体真实尺度。
- 在两个公开数据集上体积估算误差降低近30%。
- 适合精准营养、智能饮食评估等健康应用。
与饮食相关的慢性病(如肥胖、糖尿病)日益增多,准确监测食物摄入量成为迫切需求。尽管近年来人工智能辅助饮食评估取得进展,但从单目图像恢复食物尺寸(即“吃了多少”)仍面临严重挑战。现有3D重建方法虽能实现良好几何还原,却无法恢复物体的真实世界尺度,限制了其在精准营养中的应用。本文通过将3D计算机视觉与数字健康结合,提出一种从单张图像恢复真实尺度3D模型的方法。该方法利用在大规模数据集上训练的模型提取丰富视觉特征,以估计重建物体的实际尺度,从而将单视角3D重建转化为具有物理意义的真实尺寸模型。在两个公开数据集上的大量实验与消融研究显示,本方法显著优于现有技术,平均绝对体积估算误差降低近30%,展现出在精准营养领域的巨大潜力。代码已开源:https://gitlab.com/viper-purdue/size-matters
原文摘要 · Abstract (English)
The rise of chronic diseases related to diet, such as obesity and diabetes, emphasizes the need for accurate monitoring of food intake. While AI-driven dietary assessment has made strides in recent years, the ill-posed nature of recovering size (portion) information from monocular images for accurate estimation of ``how much did you eat?'' is a pressing challenge. Some 3D reconstruction methods have achieved impressive geometric reconstruction but fail to recover the crucial real-world scale of the reconstructed object, limiting its usage in precision nutrition. In this paper, we bridge the gap between 3D computer vision and digital health by proposing a method that recovers a true-to-scale 3D reconstructed object from a monocular image. Our approach leverages rich visual features extracted from models trained on large-scale datasets to estimate the scale of the reconstructed object. This learned scale enables us to convert single-view 3D reconstructions into true-to-life, physically meaningful models. Extensive experiments and ablation studies on two publicly available datasets show that our method consistently outperforms existing techniques, achieving nearly a 30% reduction in mean absolute volume-estimation error, showcasing its potential to enhance the domain of precision nutrition. Code: https://gitlab.com/viper-purdue/size-matters
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。