首个含营养信息的3D食物数据集,助力精准食量估算。
MetaFood3D: 3D Food Dataset with Nutrition Values
- 构建743个3D食物模型,覆盖131类,含详细营养与重量数据。
- 实验显示其可显著提升食量估计算法性能,填补视频与扫描数据差距。
- 适合做食物分析、营养计算和生成式3D建模的研究者使用。
食品计算在计算机视觉中既重要又具挑战性,因其广泛存在于各类应用的数据集中,涵盖分类、实例分割到3D重建。食物形态多变、纹理复杂,且包含语言描述和营养数据等多模态信息,使现代视觉算法面临严峻考验。3D食物建模因能处理任意视角并直观表示食物份量,成为新前沿。然而,现有3D数据集普遍缺乏营养值,制约了算法发展。为此,我们提出MetaFood3D,包含743个精心扫描标注的3D食物对象,覆盖131个类别,附带详尽营养信息、重量及链接至营养数据库的食物编码。该数据集强调类内多样性,提供纹理网格、RGB-D视频与分割掩码等多模态数据。实验表明,该数据集显著提升了食量估计算法性能,揭示了视频捕捉与3D扫描数据间的差距,并展现出在生成合成进食场景与3D食物方面的潜力。
原文摘要 · Abstract (English)
Food computing is both important and challenging in computer vision (CV). It significantly contributes to the development of CV algorithms due to its frequent presence in datasets across various applications, ranging from classification and instance segmentation to 3D reconstruction. The polymorphic shapes and textures of food, coupled with high variation in forms and vast multimodal information, including language descriptions and nutritional data, make food computing a complex and demanding task for modern CV algorithms. 3D food modeling is a new frontier for addressing food related problems, due to its inherent capability to deal with random camera views and its straightforward representation for calculating food portion size. However, the primary hurdle in the development of algorithms for food object analysis is the lack of nutrition values in existing 3D datasets. Moreover, in the broader field of 3D research, there is a critical need for domain-specific test datasets. To bridge the gap between general 3D vision and food computing research, we introduce MetaFood3D. This dataset consists of 743 meticulously scanned and labeled 3D food objects across 131 categories, featuring detailed nutrition information, weight, and food codes linked to a comprehensive nutrition database. Our MetaFood3D dataset emphasizes intra-class diversity and includes rich modalities such as textured mesh files, RGB-D videos, and segmentation masks. Experimental results demonstrate our dataset's strong capabilities in enhancing food portion estimation algorithms, highlight the gap between video captures and 3D scanned data, and showcase the strengths of MetaFood3D in generating synthetic eating occasion data and 3D food objects.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。