arXiv:2509.13268cs.LG2025-09被引 1

用文字记录估算饮食能量与宏量营养素,仅靠大模型微调就可实现高精度。

LLMs for energy and macronutrients estimation using only text data from 24-hour dietary recalls: a parameter-efficient fine-tuning experiment using a 10-shot prompt

  • 用链式思考提示+参数高效微调,让语言模型从食物文字描述中预测营养值。
  • 微调后能量预测误差降至190.90千卡,相关系数超0.89,显著优于原始模型。
  • 适合开发无需拍照的轻量级饮食监测工具,尤其适用于青少年群体。

背景:大多数营养估算的人工智能工具依赖图像输入。但大型语言模型(LLMs)仅凭食物消费的文字描述能否准确预测营养成分尚不明确。若有效,将实现无需照片的简化饮食监测。方法:我们使用了12-19岁青少年在国家健康与营养调查(NHANES)中的24小时膳食回顾数据。采用开源量化版LLM,通过10样本链式思考提示,仅基于列出食物及数量的文本字符串,估算能量和五种宏量营养素。随后应用参数高效微调(PEFT)评估预测精度是否提升。以NHANES计算值为能量、蛋白质、碳水化合物、总糖、膳食纤维和总脂肪的真实标签。结果:在11,281名青少年(49.9%男性,平均年龄15.4岁)的合并数据集中,原始LLM表现较差,能量均方绝对误差(MAE)为652.08,各终点林氏一致性相关系数(Lin's CCC)均低于0.46。相反,微调模型表现显著提升,能量MAE在不同子集间为171.34至190.90,所有结果的林氏一致性相关系数均超过0.89。结论:当采用链式思考提示并经由PEFT微调后,仅接收文本输入的开源LLM可准确预测24小时膳食回顾中的能量与宏量营养素。该方法有望用于低负担的文本驱动饮食监测工具。

原文摘要 · Abstract (English)

BACKGROUND: Most artificial intelligence tools used to estimate nutritional content rely on image input. However, whether large language models (LLMs) can accurately predict nutritional values based solely on text descriptions of foods consumed remains unknown. If effective, this approach could enable simpler dietary monitoring without the need for photographs. METHODS: We used 24-hour dietary recalls from adolescents aged 12-19 years in the National Health and Nutrition Examination Survey (NHANES). An open-source quantized LLM was prompted using a 10-shot, chain-of-thought approach to estimate energy and five macronutrients based solely on text strings listing foods and their quantities. We then applied parameter-efficient fine-tuning (PEFT) to evaluate whether predictive accuracy improved. NHANES-calculated values served as the ground truth for energy, proteins, carbohydrates, total sugar, dietary fiber and total fat. RESULTS: In a pooled dataset of 11,281 adolescents (49.9% male, mean age 15.4 years), the vanilla LLM yielded poor predictions. The mean absolute error (MAE) was 652.08 for energy and the Lin's CCC <0.46 across endpoints. In contrast, the fine-tuned model performed substantially better, with energy MAEs ranging from 171.34 to 190.90 across subsets, and Lin's CCC exceeding 0.89 for all outcomes. CONCLUSIONS: When prompted using a chain-of-thought approach and fine-tuned with PEFT, open-source LLMs exposed solely to text input can accurately predict energy and macronutrient values from 24-hour dietary recalls. This approach holds promise for low-burden, text-based dietary monitoring tools.

大模型饮食分析文本预测营养估算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。