用SAM模型打造食品图像半自动标注工具,让营养学家也能轻松用AI。
A SAM based Tool for Semi-Automatic Food Annotation
- 基于SAM模型,通过用户交互实现食品图像的提示式分割。
- 发布专用于食物分割的微调模型MealSAM(ViT-B骨干)。
- 适合非AI背景的营养研究者使用,推动食品数据共建。
人工智能在食品与营养研究中的发展受限于标注数据的匮乏。尽管高效的食物分割与分类模型不断涌现,但其实际应用往往需要掌握人工智能与机器学习知识,对营养科学领域的非专业人士构成挑战。为此,我们展示了一款基于Segment Anything Model(SAM)的半自动食品图像标注工具,支持用户通过交互式提示完成食物分割,可进一步对餐食图像中的食物进行分类,并在需要时标注重量或体积。此外,我们发布了经过微调的SAM掩码解码器——MealSAM,采用ViT-B骨干网络,专门优化于食物图像分割任务。本工作不仅旨在推动领域内协作与更多标注数据的积累,更致力于将先进AI技术转化为面向广大用户的实用工具。
原文摘要 · Abstract (English)
The advancement of artificial intelligence (AI) in food and nutrition research is hindered by a critical bottleneck: the lack of annotated food data. Despite the rise of highly efficient AI models designed for tasks such as food segmentation and classification, their practical application might necessitate proficiency in AI and machine learning principles, which can act as a challenge for non-AI experts in the field of nutritional sciences. Alternatively, it highlights the need to translate AI models into user-friendly tools that are accessible to all. To address this, we present a demo of a semi-automatic food image annotation tool leveraging the Segment Anything Model (SAM). The tool enables prompt-based food segmentation via user interactions, promoting user engagement and allowing them to further categorise food items within meal images and specify weight/volume if necessary. Additionally, we release a fine-tuned version of SAM's mask decoder, dubbed MealSAM, with the ViT-B backbone tailored specifically for food image segmentation. Our objective is not only to contribute to the field by encouraging participation, collaboration, and the gathering of more annotated food data but also to make AI technology available for a broader audience by translating AI into practical tools.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。