arXiv:2503.04755cs.CYcs.CL2025-03被引 1

仅用食物帖子标题估算营养成分,助力饮食行为研究

NutriTransform: Estimating Nutritional Information From Online Food Posts

  • 结合美国农业部数据库与文本嵌入技术,从标题推断营养成分
  • 在真实数据上验证有效,应用于超过50万条Reddit帖子
  • 适合关注饮食趋势、营养分析的科研与应用人员

从在线食物帖子中提取营养信息极具挑战性,尤其当用户未明确标注宏量营养素时。本文提出一种高效直接的方法,仅基于食物帖子标题估算宏量营养素。该方法融合美国农业部公开食品数据库与先进的文本嵌入技术。我们在标注的食物数据集上评估了该方法的有效性,并将其应用于Reddit热门子版块/r/food中超过50万条真实帖子,揭示了基于估算宏量营养成分的食物分享行为趋势。本工作为研究人员和实践者仅使用文本数据估算热量与营养成分提供了基础。

原文摘要 · Abstract (English)

Deriving nutritional information from online food posts is challenging, particularly when users do not explicitly log the macro-nutrients of a shared meal. In this work, we present an efficient and straightforward approach to approximating macro-nutrients based solely on the titles of food posts. Our method combines a public food database from the U.S. Department of Agriculture with advanced text embedding techniques. We evaluate the approach on a labeled food dataset, demonstrating its effectiveness, and apply it to over 500,000 real-world posts from Reddit's popular /r/food subreddit to uncover trends in food-sharing behavior based on the estimated macro-nutrient content. Altogether, this work lays a foundation for researchers and practitioners aiming to estimate caloric and nutritional content using only text data.

营养估算文本分析饮食行为

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。