用文字描述预测食物健康评分,让语言变营养信息
Semantic Nutrition Estimation: Predicting Food Healthfulness from Text Descriptions
- 融合文本嵌入与营养数据库,用多头神经网络估算营养成分
- 预测准确率中位数达R²=0.81,健康评分与真实值相关性r=0.77
- 适合做饮食评估工具,尤其适用于非专业场景的用户应用
精准的营养评估对公共健康至关重要,但现有系统依赖详细数据,常无法从日常食物描述中获取。本文提出一种机器学习流程,仅通过文本描述预测完整的食品综合评分2.0(Food Compass Score 2.0, FCS)。方法结合语义文本嵌入、词汇模式、领域启发规则与美国农业部膳食研究食品及营养数据库(FNDDS)数据,构建混合特征向量,利用多头神经网络估算生成FCS所需的营养成分。模型在个体营养素预测上达到中位R²=0.81;预测的FCS与公开值具有强相关性(皮尔逊相关系数r=0.77),平均绝对差为14.0分。尽管对模糊或加工食品误差较大,但该方法成功将自然语言转化为可操作的营养信息,支持大规模饮食评估在消费应用与研究中的落地。
原文摘要 · Abstract (English)
Accurate nutritional assessment is critical for public health, but existing profiling systems require detailed data often unavailable or inaccessible from colloquial text descriptions of food. This paper presents a machine learning pipeline that predicts the comprehensive Food Compass Score 2.0 (FCS) from text descriptions. Our approach uses multi-headed neural networks to process hybrid feature vectors that combine semantic text embeddings, lexical patterns, and domain heuristics, alongside USDA Food and Nutrient Database for Dietary Studies (FNDDS) data. The networks estimate the nutrient and food components necessary for the FCS algorithm. The system demonstratedstrong predictive power, achieving a median R^2 of 0.81 for individual nutrients. The predicted FCS correlated strongly with published values (Pearson's r = 0.77), with a mean absolute difference of 14.0 points. While errors were largest for ambiguous or processed foods, this methodology translates language into actionable nutritional information, enabling scalable dietary assessment for consumer applications and research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。