用单张照片估算中式菜肴营养,无需深度传感器。
OmniFood8K: Single-Image Nutrition Estimation via Hierarchical Frequency-Aligned Fusion

- 从单张图像生成深度图并融合多尺度频域特征
- 在8036张中式菜品图像上实现高精度营养预测
- 适合健康饮食管理与智能厨房应用
准确估计食物营养对促进健康饮食和个性化膳食管理至关重要。现有食品数据集多聚焦西式餐饮,缺乏对中式菜肴的充分覆盖,限制了中文餐食的营养估算精度。此外,许多先进方法依赖深度传感器,难以在日常场景中应用。为此,我们提出OmniFood8K,一个包含8,036个食物样本的多模态数据集,每样本配有详细营养标注和多视角图像。同时构建NutritionSynth-115K大规模合成数据集,引入组合变化并保持精确营养标签。提出端到端营养预测框架:首先从单张RGB图像预测深度图,并设计尺度-平移残差适配器(SSRA)以保证全局尺度一致性和局部结构保留;其次提出频域分层对齐融合模块(FAFM),在频域中对齐并融合RGB与深度特征;最后设计基于掩码的预测头(MPH),通过动态通道选择强调关键食材区域,提升预测准确性。在多个数据集上的大量实验表明,该方法优于现有方法。
原文摘要 · Abstract (English)
Accurate estimation of food nutrition plays a vital role in promoting healthy dietary habits and personalized diet management. Most existing food datasets primarily focus on Western cuisines and lack sufficient coverage of Chinese dishes, which restricts accurate nutritional estimation for Chinese meals. Moreover, many state-of-the-art nutrition prediction methods rely on depth sensors, restricting their applicability in daily scenarios. To address these limitations, we introduce OmniFood8K, a comprehensive multimodal dataset comprising 8,036 food samples, each with detailed nutritional annotations and multi-view images. In addition, to enhance models' capability in nutritional prediction, we construct NutritionSynth-115K, a large-scale synthetic dataset that introduces compositional variations while preserving precise nutritional labels. Moreover, we propose an end-to-end framework for nutritional prediction from a single RGB image. First, we predict a depth map from a single RGB image and design the Scale-Shift Residual Adapter (SSRA) to refine it for global scale consistency and local structural preservation. Second, we propose the Frequency-Aligned Fusion Module (FAFM) to hierarchically align and fuse RGB and depth features in the frequency domain. Finally, we design a Mask-based Prediction Head (MPH) to emphasize key ingredient regions via dynamic channel selection for more accurate prediction. Extensive experiments on multiple datasets demonstrate the superiority of our method over existing approaches. Project homepage: https://yudongjian.github.io/OmniFood8K-food/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。