用2D图像推算食物营养,无需深度传感器
PortionNet: Distilling 3D Geometric Knowledge for Food Nutrition Estimation
- 训练时学点云几何特征,推理只用普通照片
- 在MetaFood3D上体积和热量估计均达顶尖水平
- 适合手机端部署,通用性强且无需特殊硬件
从单张图像准确估算食物营养十分困难,因缺乏三维信息。尽管基于深度的方法能提供可靠几何数据,但受限于深度传感器,难以在多数智能手机上使用。为此,我们提出PortionNet,一种新型跨模态知识蒸馏框架:训练时从点云学习几何特征,推理时仅需RGB图像。该方法采用双模式训练策略,通过轻量级适配网络模拟点云表示,实现无专用硬件的伪3D推理。PortionNet在MetaFood3D数据集上达到当前最优性能,无论是体积还是能量估计均超越所有先前方法。在SimpleFood45上的跨数据集评估也表明其能量估计具备强泛化能力。
原文摘要 · Abstract (English)
Accurate food nutrition estimation from single images is challenging due to the loss of 3D information. While depth-based methods provide reliable geometry, they remain inaccessible on most smartphones because of depth-sensor requirements. To overcome this challenge, we propose PortionNet, a novel cross-modal knowledge distillation framework that learns geometric features from point clouds during training while requiring only RGB images at inference. Our approach employs a dual-mode training strategy where a lightweight adapter network mimics point cloud representations, enabling pseudo-3D reasoning without any specialized hardware requirements. PortionNet achieves state-of-the-art performance on MetaFood3D, outperforming all previous methods in both volume and energy estimation. Cross-dataset evaluation on SimpleFood45 further demonstrates strong generalization in energy estimation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。