用大模型从餐厅评论中自动识别肠胃病症状和相关食物
Review GIDE -- Restaurant Review Gastrointestinal Illness Detection and Extraction with Large Language Models
- 设计专家标注方案,用提示词让大模型同时完成疾病检测、症状与食物提取
- 大模型在三项任务上微调后均达90%以上F1值,优于小规模精调模型
- 适合公共卫生监测、数据驱动的流行病学研究者参考
食源性肠胃炎是英国常见健康问题,但多数病例未就医,传统监测难以覆盖。随着在线餐厅评论增多及大语言模型(LLMs)发展,可通过公众报告扩展疾病监测。本研究基于专家设计的标注方案,对Yelp开放数据集进行标注,涵盖症状与食物等细节信息。评估了开源大模型在肠胃病检测、症状提取、食物提取三项任务上的表现,并与专用于这些任务的RoBERTa模型对比。结果表明,采用提示工程的大模型在三项任务上微平均F1均超90%,仅靠提示即超越部分小模型。在三个偏倚实验中,大模型表现出良好鲁棒性。结果表明,利用公开评论文本与大模型可高效提取关键信息,显著提升公共卫生监测能力。尽管大模型偏差较小,但餐厅评论数据本身存在局限,需谨慎解读结果。
原文摘要 · Abstract (English)
Foodborne gastrointestinal (GI) illness is a common cause of ill health in the UK. However, many cases do not interact with the healthcare system, posing significant challenges for traditional surveillance methods. The growth of publicly available online restaurant reviews and advancements in large language models (LLMs) present potential opportunities to extend disease surveillance by identifying public reports of GI illness. In this study, we introduce a novel annotation schema, developed with experts in GI illness, applied to the Yelp Open Dataset of reviews. Our annotations extend beyond binary disease detection, to include detailed extraction of information on symptoms and foods. We evaluate the performance of open-weight LLMs across these three tasks: GI illness detection, symptom extraction, and food extraction. We compare this performance to RoBERTa-based classification models fine-tuned specifically for these tasks. Our results show that using prompt-based approaches, LLMs achieve micro-F1 scores of over 90% for all three of our tasks. Using prompting alone, we achieve micro-F1 scores that exceed those of smaller fine-tuned models. We further demonstrate the robustness of LLMs in GI illness detection across three bias-focused experiments. Our results suggest that publicly available review text and LLMs offer substantial potential for public health surveillance of GI illness by enabling highly effective extraction of key information. While LLMs appear to exhibit minimal bias in processing, the inherent limitations of restaurant review data highlight the need for cautious interpretation of results.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。