10万张食物多模态数据集,可溯源且免费公开使用
MM-Food-100K: A 100,000-Sample Multimodal Food Intelligence Dataset with Verifiable Provenance
- 基于社区贡献与AI质检构建,每条数据可追溯至钱包地址
- 在10万样本上微调大模型,营养预测准确率显著优于基线
- 适合研究食物识别、可解释性与数据确权的学者与开发者
我们提出MM-Food-100K,一个包含10万样本的多模态食物智能数据集,具备可验证的数据来源。它是原始120万张经质量筛选的食物图像语料库的约10%开放子集,涵盖菜名、起源地区等丰富标注信息。数据通过六周时间从超过8.7万名贡献者收集,采用Codatta贡献模型,结合社区众包与可配置的AI辅助质检;每条提交均通过链下安全账本绑定钱包地址以实现可追溯性,未来将上线链上协议。本文描述数据结构、处理流程与质量保障机制,并通过在图像营养预测任务上微调ChatGPT 5、ChatGPT OSS和Qwen-Max模型验证其有效性。微调后在标准指标上均取得一致提升,结果主要报告于MM-Food-100K子集。数据集公开免费获取,其余90%保留用于潜在商业用途,并向贡献者分享收益。
原文摘要 · Abstract (English)
We present MM-Food-100K, a public 100,000-sample multimodal food intelligence dataset with verifiable provenance. It is a curated approximately 10% open subset of an original 1.2 million, quality-accepted corpus of food images annotated for a wide range of information (such as dish name, region of creation). The corpus was collected over six weeks from over 87,000 contributors using the Codatta contribution model, which combines community sourcing with configurable AI-assisted quality checks; each submission is linked to a wallet address in a secure off-chain ledger for traceability, with a full on-chain protocol on the roadmap. We describe the schema, pipeline, and QA, and validate utility by fine-tuning large vision-language models (ChatGPT 5, ChatGPT OSS, Qwen-Max) on image-based nutrition prediction. Fine-tuning yields consistent gains over out-of-box baselines across standard metrics; we report results primarily on the MM-Food-100K subset. We release MM-Food-100K for publicly free access and retain approximately 90% for potential commercial access with revenue sharing to contributors.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。