提升细粒度食物图像分类准确率,助力精准营养管理
Feature-Enhanced TResNet for Fine-Grained Food Image Classification
- 引入风格重校准与深度通道注意力机制增强特征提取
- 在两个中文食物数据集上分别达81.37%和80.29%准确率
- 适合做智能膳食评估与个性化推荐系统的研究者参考
食物不仅关乎人类健康,也是文化认同与情感连接的媒介。在精准营养背景下,准确识别与分类食物图像对饮食监测、营养估算和个性化健康管理至关重要。然而,由于相似菜肴间视觉差异细微,细粒度食物分类仍具挑战。为此,我们提出特征增强型TResNet(FE-TResNet),一种专为细粒度场景优化的深度学习模型。基于TResNet架构,该模型融合风格重校准模块(StyleRM)与深度通道注意力(DCA),强化特征表达并突出食物间的细微差异。在两个基准中文食物数据集——ChineseFoodNet与CNFOOD-241上,FE-TResNet分别取得81.37%和80.29%的分类准确率。结果验证了其有效性,凸显其在精准营养系统中实现智能膳食评估与个性化推荐的关键潜力。
原文摘要 · Abstract (English)
Food is not only essential to human health but also serves as a medium for cultural identity and emotional connection. In the context of precision nutrition, accurately identifying and classifying food images is critical for dietary monitoring, nutrient estimation, and personalized health management. However, fine-grained food classification remains challenging due to the subtle visual differences among similar dishes. To address this, we propose Feature-Enhanced TResNet (FE-TResNet), a novel deep learning model designed to improve the accuracy of food image recognition in fine-grained scenarios. Built on the TResNet architecture, FE-TResNet integrates a Style-based Recalibration Module (StyleRM) and Deep Channel-wise Attention (DCA) to enhance feature extraction and emphasize subtle distinctions between food items. Evaluated on two benchmark Chinese food datasets-ChineseFoodNet and CNFOOD-241-FE-TResNet achieved high classification accuracies of 81.37% and 80.29%, respectively. These results demonstrate its effectiveness and highlight its potential as a key enabler for intelligent dietary assessment and personalized recommendations in precision nutrition systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。