用手机拍食物生成3D模型,无需标定物或深度传感器也能准估体积。
VolE: A Point-cloud Framework for Food 3D Reconstruction and Volume Estimation
- 基于手机自由运动拍摄,结合AR技术重建食物3D模型。
- 在多个数据集上达到2.22%的平均绝对百分比误差,优于现有方法。
- 无需参考物或深度信息,适合移动健康应用快速部署。
准确的食物体积估计对医疗营养管理和健康监测至关重要,但现有方法常受限于单一数据源,依赖专用硬件(如3D扫描仪)、特定传感器信息(如深度图)或需参考物进行相机标定。本文提出VolE框架,利用支持AR的移动设备驱动的3D重建来估算食物体积。该框架通过自由运动拍摄图像与相机位姿,生成高精度3D模型;同时采用无参考、无深度的视频分割策略实现食物掩码生成。我们还构建了一个包含先前基准中缺失挑战场景的新食物数据集。实验表明,VolE在多个数据集上均超越现有方法,达到2.22%的平均绝对百分比误差(MAPE),显著提升食物体积估计性能。
原文摘要 · Abstract (English)
Accurate food volume estimation is crucial for medical nutrition management and health monitoring applications, but current food volume estimation methods are often limited by mononuclear data, leveraging single-purpose hardware such as 3D scanners, gathering sensor-oriented information such as depth information, or relying on camera calibration using a reference object. In this paper, we present VolE, a novel framework that leverages mobile device-driven 3D reconstruction to estimate food volume. VolE captures images and camera locations in free motion to generate precise 3D models, thanks to AR-capable mobile devices. To achieve real-world measurement, VolE is a reference- and depth-free framework that leverages food video segmentation for food mask generation. We also introduce a new food dataset encompassing the challenging scenarios absent in the previous benchmarks. Our experiments demonstrate that VolE outperforms the existing volume estimation techniques across multiple datasets by achieving 2.22 % MAPE, highlighting its superior performance in food volume estimation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。