arXiv:2507.04364cs.CLcs.SI2025-07被引 1

对比三大模型在两类公共卫生议题中的情感识别能力,发现平台与话题影响检测准确率。

Large Language Models' Varying Accuracy in Recognizing Risk-Promoting and Health-Supporting Sentiments in Public Health Discourse: The Cases of HPV Vaccination and Heated Tobacco Products

  • 用真实社交数据+人工标注,测试GPT、Gemini、LLAMA对风险与支持性言论的分类能力。
  • 模型在Facebook上识别风险言论更准,推特上识别支持性言论更准,但中立内容难把握。
  • 提醒研究者需根据议题和平台选择模型,警惕训练数据带来的偏见。

机器学习被广泛用于分析大规模健康相关公共话语,但其在准确识别不同健康情感方面的表现仍存疑问。本文研究了三种主流大语言模型(GPT、Gemini、LLAMA)在人类乳头瘤病毒(HPV)疫苗和加热烟草产品(HTPs)两大关键公共卫生议题中,识别风险促进与健康支持性情感的准确性。基于来自Facebook和推特的数据,我们构建了支持或反对推荐健康行为的消息集,并以人工标注为金标准进行情感分类。结果表明,三类模型整体上对风险促进与健康支持性情感均表现出较高准确性,但在平台、健康议题和模型类型间存在显著差异。具体而言,模型在Facebook上对风险促进性言论的识别更准确,而在推特上对健康支持性言论的检测更为精准。此外,分析还显示模型在可靠识别中立信息方面面临挑战。这些结果强调,在开展公共卫生分析时,必须谨慎选择并验证语言模型,尤其要关注训练数据潜在偏差可能导致某些观点被高估或低估。

原文摘要 · Abstract (English)

Machine learning methods are increasingly applied to analyze health-related public discourse based on large-scale data, but questions remain regarding their ability to accurately detect different types of health sentiments. Especially, Large Language Models (LLMs) have gained attention as a powerful technology, yet their accuracy and feasibility in capturing different opinions and perspectives on health issues are largely unexplored. Thus, this research examines how accurate the three prominent LLMs (GPT, Gemini, and LLAMA) are in detecting risk-promoting versus health-supporting sentiments across two critical public health topics: Human Papillomavirus (HPV) vaccination and heated tobacco products (HTPs). Drawing on data from Facebook and Twitter, we curated multiple sets of messages supporting or opposing recommended health behaviors, supplemented with human annotations as the gold standard for sentiment classification. The findings indicate that all three LLMs generally demonstrate substantial accuracy in classifying risk-promoting and health-supporting sentiments, although notable discrepancies emerge by platform, health issue, and model type. Specifically, models often show higher accuracy for risk-promoting sentiment on Facebook, whereas health-supporting messages on Twitter are more accurately detected. An additional analysis also shows the challenges LLMs face in reliably detecting neutral messages. These results highlight the importance of carefully selecting and validating language models for public health analyses, particularly given potential biases in training data that may lead LLMs to overestimate or underestimate the prevalence of certain perspectives.

大模型情感分析公共卫生社交媒体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。