用菜名+图片比单用图片更准估算卡路里。
Multimodal ML: Quantifying the Improvement of Calorie Estimation Through Image-Text Pairs
- 结合菜名与图片的多模态模型提升估算精度。
- 平均误差从84.76降至83.70大卡,降幅1.25%。
- 适合做饮食健康应用或营养分析系统的人看。
本文评估了短文本输入(如菜品名称)对卡路里估算的提升效果,对比仅使用图像的基线模型。利用TensorFlow和Google整理的Nutrition5k数据集,训练了仅图像的CNN与接受图像和文本输入的多模态CNN。结果显示,多模态模型的卡路里估算平均绝对误差(MAE)从84.76大卡降至83.70大卡,改善1.06大卡,相对减少1.25%,且差异具有统计显著性。
原文摘要 · Abstract (English)
This paper determines the extent to which short textual inputs (in this case, names of dishes) can improve calorie estimation compared to an image-only baseline model and whether any improvements are statistically significant. Utilizes the TensorFlow library and the Nutrition5k dataset (curated by Google) to train both an image-only CNN and multimodal CNN that accepts both text and an image as input. The MAE of calorie estimations was reduced by 1.06 kcal from 84.76 kcal to 83.70 kcal (1.25% improvement) when using the multimodal model.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。