arXiv:2604.10425cs.CV2026-04ACL被引 4

构建多视角食物评估基准,推动视觉语言模型在细粒度饮食理解上的进步。

DiningBench: A Hierarchical Multi-view Benchmark for Perception and Reasoning in the Dietary Domain

  • 构建分层多视角数据集,覆盖细粒度分类、营养估算与视觉问答三类任务。
  • 3021道菜品平均5.27张图,含真实营养数据与相似菜品难样本。
  • 揭示当前模型在精细视觉区分和营养推理中的五大失败模式,适合饮食智能研究者。

近年来,视觉语言模型(VLMs)在通用视觉理解方面取得显著进展,但在食物领域仍受限于粗粒度类别、单视角图像和不准确的元数据。为此,我们提出DiningBench,一个分层的多视角基准,用于评估VLM在三个认知复杂度层次上的表现:细粒度分类、营养估算和视觉问答。不同于以往数据集,DiningBench包含3,021种不同菜品,每道菜平均5.27张图像,包含来自相同菜单的细粒度“难”负样本及基于验证的精确营养数据。我们对29个先进开源与专有模型进行了广泛评估。实验表明,尽管当前VLM在一般推理中表现良好,但在细粒度视觉区分和精确营养推理上仍存在明显短板。此外,我们系统研究了多视角输入与思维链推理的影响,识别出五种主要失败模式。DiningBench为下一代食物导向的VLM研究提供了具有挑战性的测试平台。所有代码已开源于 https://github.com/meituan/DiningBench。

原文摘要 · Abstract (English)

Recent advancements in Vision-Language Models (VLMs) have revolutionized general visual understanding. However, their application in the food domain remains constrained by benchmarks that rely on coarse-grained categories, single-view imagery, and inaccurate metadata. To bridge this gap, we introduce DiningBench, a hierarchical, multi-view benchmark designed to evaluate VLMs across three levels of cognitive complexity: Fine-Grained Classification, Nutrition Estimation, and Visual Question Answering. Unlike previous datasets, DiningBench comprises 3,021 distinct dishes with an average of 5.27 images per entry, incorporating fine-grained "hard" negatives from identical menus and rigorous, verification-based nutritional data. We conduct an extensive evaluation of 29 state-of-the-art open-source and proprietary models. Our experiments reveal that while current VLMs excel at general reasoning, they struggle significantly with fine-grained visual discrimination and precise nutritional reasoning. Furthermore, we systematically investigate the impact of multi-view inputs and Chain-of-Thought reasoning, identifying five primary failure modes. DiningBench serves as a challenging testbed to drive the next generation of food-centric VLM research. All codes are released in https://github.com/meituan/DiningBench.

视觉语言模型食物识别多视角分析营养估算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。