arXiv:2511.20652cs.HCcs.AI2025-11被引 4

首次通过真实临床试验评估大模型在营养领域的实际效果

When LLMs Can't Help: Real-World Evaluation of LLMs in Nutrition

  • 用大模型增强规则聊天机器人,实现对话多样性与营养建议
  • 7周试验中81人参与,大模型未显著改善饮食与情绪结果
  • 揭示内在测试与真实应用间巨大差距,强调以人为中心设计

大语言模型(LLM)在营养领域的信任度日益提升,但其外部评估严重不足。该领域以随机对照试验(RCT)为金标准,专家要求基于证据部署。尽管大模型在内部评估中表现良好,但此类研究多限于模拟场景。本文首次开展涉及大模型的营养领域随机对照试验。我们通过两个大模型功能增强规则型聊天机器人:(1) 消息重述以提升对话多样性与用户参与度;(2) 使用微调模型提供营养建议。在为期七周、共81名参与者的研究中,对比了含与不含大模型功能的聊天机器人版本,评估其对饮食改善、情绪健康和用户参与度的影响。尽管大模型在内部评估中表现优异,但在真实世界应用中并未带来一致的积极效果。研究揭示了内在评估与真实影响之间的关键差距,强调需要跨学科、以用户为中心的方法。

原文摘要 · Abstract (English)

The increasing trust in large language models (LLMs), especially in the form of chatbots, is often undermined by the lack of their extrinsic evaluation. This holds particularly true in nutrition, where randomised controlled trials (RCTs) are the gold standard, and experts demand them for evidence-based deployment. LLMs have shown promising results in this field, but these are limited to intrinsic setups. We address this gap by running the first RCT involving LLMs for nutrition. We augment a rule-based chatbot with two LLM-based features: (1) message rephrasing for conversational variety and engagement, and (2) nutritional counselling through a fine-tuned model. In our seven-week RCT (n=81), we compare chatbot variants with and without LLM integration. We measure effects on dietary outcome, emotional well-being, and engagement. Despite our LLM-based features performing well in intrinsic evaluation, we find that they did not yield consistent benefits in real-world deployment. These results highlight critical gaps between intrinsic evaluations and real-world impact, emphasising the need for interdisciplinary, human-centred approaches.\footnote{We provide all of our code and results at: \\ \href{https://github.com/saeshyra/diet-chatbot-trial}{https://github.com/saeshyra/diet-chatbot-trial}}

大模型评估营养生成真实世界测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。