构建多模态食物即药物评估基准,测试模型对健康状况的适配判断能力。
FAM-Bench: A Multimodal Benchmark for Condition-Aware Food-as-Medicine Reasoning

- 基于2500个专家验证实例,融合图像与配料信息进行健康适配评估
- 包含菜品适配评分与四选一比较分析两类任务,需综合营养与视觉线索
- 适用于医疗视觉语言模型的健康推理能力评测,适合医学AI研究者
食物即药物要求模型不仅识别菜品或营养成分,还需判断具体食物是否适合特定健康状况。现有食物AI基准主要评估菜品识别、食谱理解、营养估算或通用营养问答,未覆盖健康适配决策层。我们提出FAM-Bench,一个包含2500个营养专家验证实例的多模态食物即药物基准,涵盖13种饮食相关健康状况。该基准包含两项互补任务:菜品级适配评估(根据图像和配料列表判断菜品是否适合某病症)与对比菜品分析(在四个候选菜品中按条件适配度排序)。两项任务均需整合配料证据、视觉准备线索及临床营养约束,为语言与视觉-语言模型提供标准化的可落地健康推理评测平台。
原文摘要 · Abstract (English)
Food-as-Medicine requires models to reason beyond what a dish is or what nutrition it contains: they must decide whether a concrete food choice is appropriate for a specific health condition. Existing food AI benchmarks primarily evaluate dish recognition, recipe understanding, nutrient estimation, or general nutrition question answering, leaving this health-aware decision layer largely untested. We introduce FAM-Bench, a multi-modal Food-as-Medicine benchmark with 2500 nutrition-expert-verified instances across 13 diet-related health conditions. The benchmark contains two complementary tasks: dish-level suitability assessment, where models judge whether a dish is suitable for a condition from its image and ingredient list, and comparative dish analysis, where models rank four candidate dishes by condition-specific suitability. Both tasks require integrating ingredient evidence, visual preparation cues, and clinical nutrition constraints, providing a standardized testbed for grounded health-aware reasoning in language and vision-language models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。