首个融合光谱信息的食品多模态基准,助力智能分析甜度等复杂属性。
SFOOD: A Multimodal Benchmark for Comprehensive Food Attribute Analysis Beyond RGB with Spectral Insights
- 构建首个包含光谱图像的大型食品多模态数据集
- 3266类食品、2351千条数据,涵盖甜度重量等关键属性
- 揭示光谱数据对精准感知食品属性的关键作用
随着计算机视觉和大语言模型的发展,智能化已渗透至人与车领域。然而,对于产地、数量、重量、品质、甜度等丰富食品属性,现有研究仍主要集中于类别识别。这主要受限于缺乏大规模、全面的食品基准数据集。此外,甜度、重量及细粒度分类等属性仅靠RGB相机难以准确感知。为填补这一空白并推动智能食品分析发展,本文构建了首个大规模光谱食品(SFOOD)基准套件。我们投入大量人力与设备,整合现有食品数据集,并采集数百种食品的高光谱图像,通过仪器实验测定如甜度、重量等属性。最终数据集包含3,266个食品类别和2,351千条数据点,覆盖17个主要食品类别。大量评估表明:(i) 大规模模型在食品数字化方面表现仍差,相较人与车,食品正成为最难分析的对象之一;(ii) 光谱数据对分析甜度等食品特性至关重要。该基准将开源并持续迭代,服务于各类食品分析任务。
原文摘要 · Abstract (English)
With the rise and development of computer vision and LLMs, intelligence is everywhere, especially for people and cars. However, for tremendous food attributes (such as origin, quantity, weight, quality, sweetness, etc.), existing research still mainly focuses on the study of categories. The reason is the lack of a large and comprehensive benchmark for food. Besides, many food attributes (such as sweetness, weight, and fine-grained categories) are challenging to accurately percept solely through RGB cameras. To fulfill this gap and promote the development of intelligent food analysis, in this paper, we built the first large-scale spectral food (SFOOD) benchmark suite. We spent a lot of manpower and equipment costs to organize existing food datasets and collect hyperspectral images of hundreds of foods, and we used instruments to experimentally determine food attributes such as sweetness and weight. The resulting benchmark consists of 3,266 food categories and 2,351 k data points for 17 main food categories. Extensive evaluations find that: (i) Large-scale models are still poor at digitizing food. Compared to people and cars, food has gradually become one of the most difficult objects to study; (ii) Spectrum data are crucial for analyzing food properties (such as sweetness). Our benchmark will be open source and continuously iterated for different food analysis tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。