arXiv:2606.08948cs.CVcs.AI2026-06

用饮食回忆生成合成数据,训练出首个全面分析微量营养素的视觉语言模型。

NutriMLLM: Multimodal Large Language Models for Dietary Micronutrient Analysis

论文配图:NutriMLLM: Multimodal Large Language Models for Dietary Micronutrient Analysis
图 1 · 摘自论文原文
  • 利用十年饮食回忆数据生成110万张食物图像-描述-营养三元组作为训练集。
  • 模型在真实食物图像上对65种微量营养素覆盖率达近100%,性能超商用模型。
  • 适合临床营养、个性化饮食指导及大规模人群营养监测场景。

从食物图像中全面估计微量营养素可提升临床营养护理水平,但训练此类模型需大量多模态数据。我们发现现有主流多模态大模型(包括领先专有模型)在该任务中表现不可靠:在五个模型家族和四个独立评估基准(ASA24、SNAPMe、FNDDS、NutriBench)中,模型常回避回答或输出统计上不合理的数值。为避免高昂的人工标注成本,我们利用十年间大规模24小时饮食回忆数据,将其转化为结构化提示用于文本到图像生成。该流程构建了约110万条图像-描述-营养三元组,每条包含一张生成的食物图像及完整的65种微量营养素标签。据我们所知,这是首个计划公开发布的大型合成食物图像多营养注释数据集。基于此数据集,微调Qwen3-VL(2B/4B/8B/30B)与GLM-4.6V-Flash,得到NutriMLLM,首个专注于全面饮食微量营养素估计的视觉语言模型家族。我们采用四维度评估框架,分别衡量回避率、幻觉率、整体可用性及单营养素数值准确性。在真实食物图像上,所有NutriMLLM变体均实现65种营养素的近全覆盖,最大版本在多数营养素上精度匹配或超越商用基线(GPT-5、Gemini 3、Claude Sonnet 4.5)。结果表明,基于回忆数据的合成监督可使图像驱动的全面微量营养素估计成为可行工程问题,支持饮食评估、个性化营养建议与人群级微量营养素监测。

原文摘要 · Abstract (English)

Comprehensive estimation of dietary micronutrients from food images could improve clinical nutrition care, but training such models requires large multimodal datasets linking diverse foods to complete nutrient profiles. We first show that existing multimodal large language models (MLLMs), including leading proprietary models, are unreliable for this task. Across five model families and four independent evaluation benchmarks (ASA24, SNAPMe, FNDDS, and NutriBench), models frequently abstained or returned statistically implausible values. To address this gap without costly expert annotation, we repurposed a decade of population-scale 24-hour dietary recalls as structured prompts for text-to-image generation. This pipeline produced a synthetic corpus of about 1.1 million image-description-nutrient triplets, each pairing a generated food image with a complete 65-nutrient label. To our knowledge, this is the largest synthetic food-image corpus with comprehensive micronutrient annotation planned for public release upon publication. Fine-tuning Qwen3-VL (2B/4B/8B/30B) and GLM-4.6V-Flash on this corpus yielded NutriMLLM, the first family of vision-language models specialized for comprehensive dietary micronutrient estimation. We evaluate these models with a four-component framework that separately measures abstention, hallucination, overall usability, and per-nutrient numerical accuracy. On real food images, every NutriMLLM variant achieved near-complete coverage across all 65 nutrients, and the largest variant matched or exceeded proprietary baselines (GPT-5, Gemini 3, and Claude Sonnet 4.5) in accuracy on most nutrients. These results show that recall-driven synthetic supervision can make image-based comprehensive micronutrient estimation a tractable engineering problem and support dietary assessment, personalized nutrition guidance, and population-scale micronutrient surveillance.

多模态营养分析合成数据视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。