arXiv:2512.14574cs.CVcs.MM2025-12

构建真实用餐场景的食品图像数据集,提升饮食管理模型实用性。

FoodLogAthl-218: Constructing a Real-World Food Image Dataset Using Dietary Management Applications

  • 从用户真实餐食记录中收集图片,按实际使用场景标注类别。
  • 包含6925张图、218类食物,14349个边界框,具自然分布与多样性。
  • 支持时序微调和上下文感知分类,适合研究真实饮食应用。

食品图像分类模型对饮食管理应用至关重要,可减轻手动记录饮食的负担。然而,现有公开数据集多依赖网络爬取图像,与用户真实用餐照片差异较大。本文提出FoodLogAthl-218,一个基于饮食管理应用FoodLog Athl收集的真实餐食图像数据集。数据集包含6,925张图像,覆盖218种食物类别,共14,349个边界框,每张图像附带用餐时间、匿名用户ID及餐食上下文等丰富元数据。不同于传统数据集先定类别再搜图的方式,本数据集以用户上传的原始照片为基础,事后标注标签,因此具有更高的类内多样性、更自然的餐食类型分布,以及非正式、未经修饰的个人使用图像。除标准分类基准外,我们还引入两个针对食品日志特性的任务:(1)遵循用户记录时间流的增量微调协议;(2)上下文感知分类任务,即单张图像含多个菜品,需结合整餐上下文分别识别各菜品。我们使用大型多模态模型评估这些任务。数据集已公开于https://huggingface.co/datasets/FoodLog/FoodLogAthl-218。

原文摘要 · Abstract (English)

Food image classification models are crucial for dietary management applications because they reduce the burden of manual meal logging. However, most publicly available datasets for training such models rely on web-crawled images, which often differ from users' real-world meal photos. In this work, we present FoodLogAthl-218, a food image dataset constructed from real-world meal records collected through the dietary management application FoodLog Athl. The dataset contains 6,925 images across 218 food categories, with a total of 14,349 bounding boxes. Rich metadata, including meal date and time, anonymized user IDs, and meal-level context, accompany each image. Unlike conventional datasets-where a predefined class set guides web-based image collection-our data begins with user-submitted photos, and labels are applied afterward. This yields greater intra-class diversity, a natural frequency distribution of meal types, and casual, unfiltered images intended for personal use rather than public sharing. In addition to (1) a standard classification benchmark, we introduce two FoodLog-specific tasks: (2) an incremental fine-tuning protocol that follows the temporal stream of users' logs, and (3) a context-aware classification task where each image contains multiple dishes, and the model must classify each dish by leveraging the overall meal context. We evaluate these tasks using large multimodal models (LMMs). The dataset is publicly available at https://huggingface.co/datasets/FoodLog/FoodLogAthl-218.

食品图像饮食管理真实数据多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。