评测大模型在新冠科学声明检测中的表现,探索其自动辟谣潜力。
Evaluating the Performance of Large Language Models in Scientific Claim Detection and Classification
- 用大模型直接判断社交媒体上的新冠声明真伪,免去繁琐训练。
- 模型对科学声明的分类准确率显著高于传统方法。
- 适合公共健康传播与信息治理研究者参考。
新冠疫情期间,社交媒体的广泛传播既促进了信息交流,也加速了虚假信息的扩散,形成‘数字谣言暴发’。这凸显了开发自动化工具以识别和传播真实信息的紧迫性。本研究评估了大型语言模型(LLMs)作为应对社交媒体上虚假信息的创新解决方案的有效性。以OpenAI的GPT和Meta的LLaMA为代表的LLMs,具备预训练、可迁移的优势,避免了传统机器学习模型训练成本高、易过拟合的问题。我们基于专门构建的数据集,评估了这些模型在检测和分类新冠相关科学声明方面的性能,旨在支持公众决策。结果表明,LLMs在该任务中展现出巨大潜力,可作为自动事实核查工具。尽管该领域尚处于起步阶段,仍需深入研究,但我们提出了一个比较分析框架,为公共卫生传播中的模型应用提供了实践路径。
原文摘要 · Abstract (English)
The pervasive influence of social media during the COVID-19 pandemic has been a double-edged sword, enhancing communication while simultaneously propagating misinformation. This \textit{Digital Infodemic} has highlighted the urgent need for automated tools capable of discerning and disseminating factual content. This study evaluates the efficacy of Large Language Models (LLMs) as innovative solutions for mitigating misinformation on platforms like Twitter. LLMs, such as OpenAI's GPT and Meta's LLaMA, offer a pre-trained, adaptable approach that bypasses the extensive training and overfitting issues associated with traditional machine learning models. We assess the performance of LLMs in detecting and classifying COVID-19-related scientific claims, thus facilitating informed decision-making. Our findings indicate that LLMs have significant potential as automated fact-checking tools, though research in this domain is nascent and further exploration is required. We present a comparative analysis of LLMs' performance using a specialized dataset and propose a framework for their application in public health communication.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。