arXiv:2607.08423cs.AI2026-07被引 1

评测视觉语言模型在营养推理与个性化健康建议中的表现

OmniFood-Bench: Evaluating VLMs for Nutrient Reasoning and Personalized Health Advice

论文配图:OmniFood-Bench: Evaluating VLMs for Nutrient Reasoning and Personalized Health Advice
图 1 · 摘自论文原文
  • 构建三阶段评估体系,涵盖感知、量化和安全建议能力
  • 模型对菜品识别准确率高,但质量估算错误率超60%
  • 适合研究医疗AI可靠性及饮食健康管理系统的开发者

大型视觉语言模型(VLMs)正被广泛应用于个性化医疗与饮食管理。然而,在食物领域,自主代理面临“系统性信息不对称”问题:视觉外观与内在营养成分之间存在巨大鸿沟。现有基准多聚焦粗粒度分类任务,难以评估真实饮食管理所需的复杂推理链——从识别隐藏成分、估算物理质量,到生成关键医疗建议。本文提出OmniFood-Bench,基于MM-Food-100K数据集构建的综合性评测基准。该基准涵盖三个渐进能力:基础感知(成分与烹饪方式)、定量推理(份量与营养分析)、安全关键建议(疾病特异性推荐)。我们评估了六种先进VLMs,包括gpt-5.1、gemini-3-flash和qwen3-vl-8B。实验揭示显著的“语义-物理鸿沟”:模型在菜品命名上接近人类水平,但在质量估计上失败率超过60%,且常为高风险糖尿病患者生成错误的安全建议。本工作确立了公共健康应用中自主代理可信性的严格标准。代码与数据集已公开:https://anonymous.4open.science/r/OmniFood-Bench-7D0B

原文摘要 · Abstract (English)

The rapid integration of Large Vision-Language Models (VLMs) into critical infrastructure promises to revolutionize personalized healthcare and dietary management. However, in the domain of food systems, autonomous agents face a unique and persistent challenge: the "Systemic Information Asymmetry" between visual appearance and intrinsic nutritional composition. Existing benchmarks primarily focus on coarse-grained classification tasks, such as food category recognition, which fail to evaluate the intricate reasoning chain required for real-world dietary management -- specifically, the ability to traverse from identifying hidden ingredients to estimating physical mass, and finally synthesizing safety-critical medical advice. In this paper, we introduce OmniFood-Bench, a comprehensive benchmark constructed from the MM-Food-100K dataset. Unlike previous works, OmniFood-Bench evaluates VLMs across three progressive capabilities: Basic Perception (Ingredients & Cooking Methods), Quantitative Reasoning (Portion Size & Nutritional Profiling), and Safety-Critical Advisory (Disease-Specific Recommendations). We evaluate six state-of-the-art VLMs, including gpt-5.1, gemini-3-flash, and qwen3-vl-8B. Our extensive experiments reveal a startling "Semantic-Physical Gap": while models achieve near-human accuracy in naming dishes, they exhibit catastrophic failure in mass estimation and frequently hallucinate benign advice for high-risk diabetic profiles. This work establishes a rigorous standard for trustworthiness in autonomous agents deployed for public health. The code and datasets are available in: https://anonymous.4open.science/r/OmniFood-Bench-7D0B

视觉语言模型营养推理健康建议评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。