无需训练的本地化饮食估计算法,提升份量估计准确率30%以上。
Open-KNEAD: Knowledge-grounded Nutrition Estimation via Agentic Decomposition

- 通过智能拆解食物并关联营养数据库,实现可追溯的逐项记录。
- 在三种菜系上优于传统检索和直接估算,最高提升53%准确率。
- 适合临床场景使用,支持本地部署且无需用户标注图像。
多模态大模型在从餐食图像进行饮食评估中应用日益广泛,此前研究表明检索增强能提升营养估算精度。但当前发现,现代多模态大模型的直接估算已达到或超过完整检索流程效果。这引发新问题:若检索不再提升整体估算,是否仍能提供临床所需的精确份量与可审计的逐项记录?我们在此框架下保持临床采纳关键要素——仅需一张未标注的餐食图像、可解释性(可审计记录)、隐私保护(本地推理)。提出 Open-KNEAD,一种无需训练、可在本地部署的知识增强代理框架。每个分解食物项通过选择性、营养感知检索对齐至食品与营养研究数据库(FNDDS)编码,生成可审计的单项记录。在两个开源多模态大模型家族及三种菜系上,Open-KNEAD 在多数主干-数据集组合中优于以往接地方法与直接估算。内部配方先验步骤进一步恢复非美式菜系中因烹饪添加的能量偏差。优势在经营养师验证的 ACETADA 数据集上最显著,本地开源代理相比两个前沿闭源模型的直接份量估算分别提升约30%和53%,且所有图像均保留在本地硬件上。我们发布 Open-KNEAD 框架及其就绪的 FNDDS 知识库。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) are increasingly used for dietary assessment from meal images, where retrieval-augmented grounding was shown to sharpen nutrition estimates. However, we find this premise no longer holds for current MLLMs. A modern MLLM's direct estimate now matches or surpasses the full retrieval pipeline. This raises a question: if retrieval no longer improves the overall estimate, can it still deliver the two things clinicians value, accurate portions and a traceable, item-by-item record? We pursue this while preserving what matters for clinical adoption: minimal user burden (a single, unannotated meal image), explainability (an auditable record), and privacy (locally hosted inference). We introduce Open-KNEAD, a knowledge-grounded agentic framework for meal nutrition estimation that is training-free and locally deployable. Each decomposed food item is grounded to a Food and Nutrient Database for Dietary Studies (FNDDS) code via selective, nutrient-aware retrieval, composing an auditable per-item record. Across two open MLLM families and three cuisines, Open-KNEAD improves portion estimates over both prior grounding methods and direct estimation in most backbone-dataset settings. An agent-internal recipe-prior step further recovers the invisible cooking-added energy that biases estimates on non-US cuisine. The advantage is largest on the dietitian-verified ACETADA dataset, where the local open agent surpasses the direct portion estimates of two frontier closed models by roughly $30\%$ and $53\%$, all while keeping every meal image on local hardware. We release the Open-KNEAD framework and its agent-ready FNDDS knowledge base.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。