构建首个覆盖80类印度菜的131万张图像数据集,助力食物识别与应用开发。
Khana: A Comprehensive Indian Cuisine Dataset
- 建立印度菜系分类体系,收录80类菜品图像
- 含131万张500x500像素图像,支持分类/分割/检索任务
- 为研究者与开发者提供真实场景应用基准
随着全球对多元饮食体验的兴趣增长,食物图像模型在提升食物识别、食谱推荐、膳食追踪和自动餐食规划等应用中日益重要。尽管已有大量食物数据集,但印度菜因地域多样性、复杂烹饪方式及缺乏全面标注数据,仍存在明显空白。本文提出Khana,一个用于印度菜图像分类、分割和检索的新基准数据集。该数据集通过构建印度菜系分类体系,包含约131,000张图像,覆盖80个类别,每张图像分辨率为500x500像素。本文详述数据集构建过程,并评估了先进模型在分类、分割和检索任务上的表现,为研究与开发提供综合性挑战性基准,推动相关技术落地应用。
原文摘要 · Abstract (English)
As global interest in diverse culinary experiences grows, food image models are essential for improving food-related applications by enabling accurate food recognition, recipe suggestions, dietary tracking, and automated meal planning. Despite the abundance of food datasets, a noticeable gap remains in capturing the nuances of Indian cuisine due to its vast regional diversity, complex preparations, and the lack of comprehensive labeled datasets that cover its full breadth. Through this exploration, we uncover Khana, a new benchmark dataset for food image classification, segmentation, and retrieval of dishes from Indian cuisine. Khana fills the gap by establishing a taxonomy of Indian cuisine and offering around 131K images in the dataset spread across 80 labels, each with a resolution of 500x500 pixels. This paper describes the dataset creation process and evaluates state-of-the-art models on classification, segmentation, and retrieval as baselines. Khana bridges the gap between research and development by providing a comprehensive and challenging benchmark for researchers while also serving as a valuable resource for developers creating real-world applications that leverage the rich tapestry of Indian cuisine. Webpage: https://khana.omkar.xyz
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。