梳理新冠疫苗立场分析中的误解,揭示方法缺陷对防疫认知的误导
Clarifying Misconceptions in COVID-19 Vaccine Sentiment and Stance Analysis and Their Implications for Vaccine Hesitancy Mitigation: A Systematic Review
- 系统回顾2020-2023年推特数据研究,按五维度分类分析方法
- 发现多数研究存在测量偏差,影响对疫苗犹豫的真实判断
- 提醒研究者规范NLP报告,避免误导公众防疫决策
机器学习进步使自然语言处理(NLP)在社交媒体中识别疫苗犹豫成为可能。大量研究显示,新冠疫苗犹豫在推特等平台持续存在。本研究注册于PROSPERO国际系统综述注册库,检索了2020年1月1日至2023年12月31日间,使用监督学习通过推特(现称X)上的立场检测或情感分析评估新冠疫苗犹豫的研究。依据五维分类体系——推文样本选取方式、自报研究类型、分类范式、标注代码本定义及结果解释进行归类。分析表明,使用立场检测与情感分析的研究在衡量疫苗犹豫时存在广泛测量偏差,报告错误严重到阻碍对个体是否拒绝接种SARS-CoV-2疫苗的准确理解。改进NLP方法报告质量,是填补疫苗犹豫话语研究知识空白的关键。
原文摘要 · Abstract (English)
Background Advances in machine learning (ML) models have increased the capability of researchers to detect vaccine hesitancy in social media using Natural Language Processing (NLP). A considerable volume of research has identified the persistence of COVID-19 vaccine hesitancy in discourse shared on various social media platforms. Methods Our objective in this study was to conduct a systematic review of research employing sentiment analysis or stance detection to study discourse towards COVID-19 vaccines and vaccination spread on Twitter (officially known as X since 2023). Following registration in the PROSPERO international registry of systematic reviews, we searched papers published from 1 January 2020 to 31 December 2023 that used supervised machine learning to assess COVID-19 vaccine hesitancy through stance detection or sentiment analysis on Twitter. We categorized the studies according to a taxonomy of five dimensions: tweet sample selection approach, self-reported study type, classification typology, annotation codebook definitions, and interpretation of results. We analyzed if studies using stance detection report different hesitancy trends than those using sentiment analysis by examining how COVID-19 vaccine hesitancy is measured, and whether efforts were made to avoid measurement bias. Results Our review found that measurement bias is widely prevalent in studies employing supervised machine learning to analyze sentiment and stance toward COVID-19 vaccines and vaccination. The reporting errors are sufficiently serious that they hinder the generalisability and interpretation of these studies to understanding whether individual opinions communicate reluctance to vaccinate against SARS-CoV-2. Conclusion Improving the reporting of NLP methods is crucial to addressing knowledge gaps in vaccine hesitancy discourse.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。