用大模型在健康社区中实现专家级情感分析,无需大量训练数据。
The Promise of Large Language Models in Digital Health: Evidence from Sentiment Analysis in Online Health Communities
- 通过提示工程整合医学专家知识,让大模型理解医疗语境中的情绪
- 在400条标注数据上,大模型表现媲美专家,与人工标注差异不显著
- 适合需要实时分析患者情绪的数字健康研究与临床干预评估
数字健康分析面临严峻挑战:患者生成的健康内容包含复杂情绪与医学语境,需专业领域知识,而传统机器学习受限于数据不足和隐私问题。在线健康社区(OHCs)中的帖子情绪混杂、术语密集、情感隐含,对情感分析(SA)提出高要求。本研究探索大语言模型(LLMs)通过上下文学习融合专家知识,实现高效精准的情感分析。我们构建结构化编码手册,系统化编码专家判断准则,使大模型通过针对性提示应用领域知识,而非依赖大规模训练。六款GPT模型与DeepSeek、LLaMA 3.1对比多种预训练模型(BioBERT变体)及词典方法,使用来自两个OHC的400条专家标注数据进行评估。结果表明,大模型性能优异,且与专家间一致性无统计学差异,说明其知识整合超越表面模式识别。不同大模型均表现出稳定性能,证明基于上下文学习的方案具备可扩展性,为数字健康分析提供突破专家资源短缺的可行路径,支持实时患者监测、干预评估与循证策略制定。
原文摘要 · Abstract (English)
Digital health analytics face critical challenges nowadays. The sophisticated analysis of patient-generated health content, which contains complex emotional and medical contexts, requires scarce domain expertise, while traditional ML approaches are constrained by data shortage and privacy limitations in healthcare settings. Online Health Communities (OHCs) exemplify these challenges with mixed-sentiment posts, clinical terminology, and implicit emotional expressions that demand specialised knowledge for accurate Sentiment Analysis (SA). To address these challenges, this study explores how Large Language Models (LLMs) can integrate expert knowledge through in-context learning for SA, providing a scalable solution for sophisticated health data analysis. Specifically, we develop a structured codebook that systematically encodes expert interpretation guidelines, enabling LLMs to apply domain-specific knowledge through targeted prompting rather than extensive training. Six GPT models validated alongside DeepSeek and LLaMA 3.1 are compared with pre-trained language models (BioBERT variants) and lexicon-based methods, using 400 expert-annotated posts from two OHCs. LLMs achieve superior performance while demonstrating expert-level agreement. This high agreement, with no statistically significant difference from inter-expert agreement levels, suggests knowledge integration beyond surface-level pattern recognition. The consistent performance across diverse LLM models, supported by in-context learning, offers a promising solution for digital health analytics. This approach addresses the critical challenge of expert knowledge shortage in digital health research, enabling real-time, expert-quality analysis for patient monitoring, intervention assessment, and evidence-based health strategies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。