arXiv:2508.09966cs.CVcs.AI2025-08被引 2

构建公开食物分析基准数据集,提升营养自动分析评估标准。

January Food Benchmark (JFB): A Public Benchmark Dataset and Evaluation Suite for Multimodal Food Analysis

  • 构建1000张真实食物图像的标注数据集JFB,支持多模态分析。
  • 提出综合评估框架,专用模型整体得分86.2,领先通用模型12.1分。
  • 适合营养分析、视觉语言模型及多模态应用研究者使用。

AI在自动化营养分析领域的进展因缺乏标准化评估方法和高质量真实世界基准数据集而受限。为此,本文提出三项主要贡献:首先,发布公开可用的January Food Benchmark(JFB)数据集,包含1000张食物图像及人工验证标注;其次,设计一套完整的基准评估框架,包括稳健的评估指标与面向应用的整体评分;第三,提供通用视觉语言模型(VLMs)与自研专用模型january/food-vision-v1的基线结果。评估显示,专用模型取得86.2的整体得分,比表现最佳的通用配置高出12.1分。本工作为研究社区提供了有价值的评估数据集与严谨的基准框架,以推动自动化营养分析的后续发展。

原文摘要 · Abstract (English)

Progress in AI for automated nutritional analysis is critically hampered by the lack of standardized evaluation methodologies and high-quality, real-world benchmark datasets. To address this, we introduce three primary contributions. First, we present the January Food Benchmark (JFB), a publicly available collection of 1,000 food images with human-validated annotations. Second, we detail a comprehensive benchmarking framework, including robust metrics and a novel, application-oriented overall score designed to assess model performance holistically. Third, we provide baseline results from both general-purpose Vision-Language Models (VLMs) and our own specialized model, january/food-vision-v1. Our evaluation demonstrates that the specialized model achieves an Overall Score of 86.2, a 12.1-point improvement over the best-performing general-purpose configuration. This work offers the research community a valuable new evaluation dataset and a rigorous framework to guide and benchmark future developments in automated nutritional analysis.

食物分析多模态基准测试营养评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。