用位置时间等上下文信息提升大模型营养分析准确率
Comprehensive Evaluation of Large Multimodal Models for Nutrition Analysis: A New Benchmark Enriched with Contextual Metadata
- 引入位置、时间等上下文元数据增强图像营养分析
- 结合元数据后模型误差降低,部分指标下降超30%
- 适合做健康科技、智能饮食应用的研究者参考
大型多模态模型(LMMs)在餐食图像营养分析中应用日益广泛,但现有研究多聚焦于封闭模型如GPT-4,对开放模型的探索不足。同时,上下文元数据与推理增强方法的交互机制仍不明确。本文研究了从地理坐标(转换为地点/场所类型)、时间戳(转化为餐次/日期类型)及食物内容中提取的上下文元数据如何提升对热量、宏量营养素(蛋白质、碳水化合物、脂肪)和份量的估计精度。我们提出了新的公开数据集ACETADA,包含经注册营养师验证的营养信息。在8个LMMs(4个开源+4个闭源)上评估发现,引入上下文元数据比仅用图像提示显著提升性能;进一步证明该信息能有效增强Chain-of-Thought、Multimodal Chain-of-Thought、Scale Hint、Few-Shot及Expert Persona等推理策略的效果。实证结果表明,合理整合元数据可显著降低预测值的平均绝对误差(MAE)和平均绝对百分比误差(MAPE),凸显上下文感知型LMM在营养分析中的潜力。
原文摘要 · Abstract (English)
Large Multimodal Models (LMMs) are increasingly applied to meal images for nutrition analysis. However, existing work primarily evaluates proprietary models, such as GPT-4. This leaves the broad range of LLMs underexplored. Additionally, the influence of integrating contextual metadata and its interaction with various reasoning modifiers remains largely uncharted. This work investigates how interpreting contextual metadata derived from GPS coordinates (converted to location/venue type), timestamps (transformed into meal/day type), and the food items present can enhance LMM performance in estimating key nutritional values. These values include calories, macronutrients (protein, carbohydrates, fat), and portion sizes. We also introduce \textbf{ACETADA}, a new food-image dataset slated for public release. This open dataset provides nutrition information verified by the dietitian and serves as the foundation for our analysis. Our evaluation across eight LMMs (four open-weight and four closed-weight) first establishes the benefit of contextual metadata integration over straightforward prompting with images alone. We then demonstrate how this incorporation of contextual information enhances the efficacy of reasoning modifiers, such as Chain-of-Thought, Multimodal Chain-of-Thought, Scale Hint, Few-Shot, and Expert Persona. Empirical results show that integrating metadata intelligently, when applied through straightforward prompting strategies, can significantly reduce the Mean Absolute Error (MAE) and Mean Absolute Percentage Error (MAPE) in predicted nutritional values. This work highlights the potential of context-aware LMMs for improved nutrition analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。