arXiv:2504.06925cs.CVcs.AI2025-04CVPR被引 17

评测6大视觉语言模型在食物识别中的表现,发现闭源模型更优但细粒度识别仍不足。

Are Vision-Language Models Ready for Dietary Assessment? Exploring the Next Frontier in AI-Powered Food Image Recognition

  • 引入多层级食物数据库与专家加权召回率评估新指标
  • 闭源模型单品类识别超90%准确率,开源自闭差距明显
  • 适合做饮食评估的AI研究者和医疗健康领域应用开发者

基于食物图像的自动膳食评估仍具挑战,需精准检测、分割与分类。视觉语言模型(VLMs)通过融合视觉与文本推理提供了新可能。本研究评估了六种前沿VLMs(ChatGPT、Gemini、Claude、Moondream、DeepSeek、LLaVA)在不同层级食物识别中的能力。为此,我们构建了FoodNExTDB——一个包含9,263张专家标注图像的数据库,涵盖10个类别(如“蛋白质来源”)、62个子类别(如“禽类”)和9种烹饪方式(如“烤制”)。共生成50,000条营养标签,由七位专家人工标注。提出新评估指标专家加权召回率(EWR),考虑标注者间差异。结果表明,闭源模型优于开源模型,在仅含单一食物的图像中实现超过90%的EWR。然而,当前VLMs在精细识别上仍面临困难,尤其在区分烹饪方式差异及外观相似食物方面,限制其在自动膳食评估中的可靠性。FoodNExTDB已公开于https://github.com/AI4Food/FoodNExtDB。

原文摘要 · Abstract (English)

Automatic dietary assessment based on food images remains a challenge, requiring precise food detection, segmentation, and classification. Vision-Language Models (VLMs) offer new possibilities by integrating visual and textual reasoning. In this study, we evaluate six state-of-the-art VLMs (ChatGPT, Gemini, Claude, Moondream, DeepSeek, and LLaVA), analyzing their capabilities in food recognition at different levels. For the experimental framework, we introduce the FoodNExTDB, a unique food image database that contains 9,263 expert-labeled images across 10 categories (e.g., "protein source"), 62 subcategories (e.g., "poultry"), and 9 cooking styles (e.g., "grilled"). In total, FoodNExTDB includes 50k nutritional labels generated by seven experts who manually annotated all images in the database. Also, we propose a novel evaluation metric, Expert-Weighted Recall (EWR), that accounts for the inter-annotator variability. Results show that closed-source models outperform open-source ones, achieving over 90% EWR in recognizing food products in images containing a single product. Despite their potential, current VLMs face challenges in fine-grained food recognition, particularly in distinguishing subtle differences in cooking styles and visually similar food items, which limits their reliability for automatic dietary assessment. The FoodNExTDB database is publicly available at https://github.com/AI4Food/FoodNExtDB.

食物识别视觉语言模型饮食评估多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。