arXiv:2412.17217q-bio.BMcs.LG2024-12被引 8

用营养成分和文本信息预测食品加工程度,准确率超92%。

Machine learning and natural language processing models to predict the extent of food processing

  • 结合营养成分与文本数据,构建机器学习模型预测食品加工等级。
  • 102种营养成分下模型F1达0.9411,13种关键营养成分下仍保持0.9284。
  • 提供在线工具,用户可输入食品营养数据即时获取加工等级预测。

超加工食品消费激增与多种不良健康效应相关。为应对由此带来的公共健康问题,我们构建了多种机器学习、深度学习及自然语言处理模型,基于食品产品营养成分数据与报告的NOVA加工等级,预测其加工程度。初始采用102项营养特征,经粗化后分别降至65项和13项(依据美国食品药品管理局标准)。在102项特征下,轻量级梯度提升机(LGBM)分类器表现最佳,F1分数为0.9411,马修斯相关系数(MCC)为0.8691;65项特征下,随机森林最优,F1为0.9345,MCC为0.8543;13项营养成分下,梯度提升模型取得最高F1(0.9284),MCC为0.8425。此外,基于NLP的模型也达到领先性能。研究还识别出影响模型表现的关键营养成分,并发布了一个交互式网页服务器,支持用户通过输入食品营养成分预测其加工等级:https://cosylab.iiitd.edu.in/food-processing/。

原文摘要 · Abstract (English)

The dramatic increase in consumption of ultra-processed food has been associated with numerous adverse health effects. Given the public health consequences linked to ultra-processed food consumption, it is highly relevant to build computational models to predict the processing of food products. We created a range of machine learning, deep learning, and NLP models to predict the extent of food processing by integrating the FNDDS dataset of food products and their nutrient profiles with their reported NOVA processing level. Starting with the full nutritional panel of 102 features, we further implemented coarse-graining of features to 65 and 13 nutrients by dropping flavonoids and then by considering the 13-nutrient panel of FDA, respectively. LGBM Classifier and Random Forest emerged as the best model for 102 and 65 nutrients, respectively, with an F1-score of 0.9411 and 0.9345 and MCC of 0.8691 and 0.8543. For the 13-nutrient panel, Gradient Boost achieved the best F1-score of 0.9284 and MCC of 0.8425. We also implemented NLP based models, which exhibited state-of-the-art performance. Besides distilling nutrients critical for model performance, we present a user-friendly web server for predicting processing level based on the nutrient panel of a food product: https://cosylab.iiitd.edu.in/food-processing/.

食品加工机器学习营养分析预测模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。