arXiv:2608.15691cs.CLcs.LG2026-08

用语义+传播力双指标,从推特抓出高危谣言集群。

BERTopic-Virality Prioritisation: A Scalable Framework for Thematic and Comparative Analysis of COVID-19 and Monkeypox Misinformation on Twitter

论文配图:BERTopic-Virality Prioritisation: A Scalable Framework for Thematic and Comparative Analysis of COVID-19 and Monkeypox Misinformation on Twitter
图 1 · 摘自论文原文
  • 先用BERTopic聚类话题,再按传播潜力排序高危主题。
  • 在三个数据集上误报检测F1达0.950,识别出前1%高传播谣言。
  • 适合公共卫生监测者快速发现隐蔽但危险的谣言线索。

疫情期间健康谣言迅速扩散,形成与公共健康指引竞争的有害叙事。现有主题建模方法将互动量视为外部结果,难以优先识别语义连贯且传播快的话题。本文提出BERTopic-VP框架,结合基于上下文嵌入的聚类(BERTopic)与事后传播力优先层(VP)。该流程还包含两级混合谣言检测模块:融合监督内容分类器与来自公共健康知识库的外部验证信号。应用于三个基准数据集(COVID-19_FNIR、Monkeypox、Constraint),框架在误报检测上达到最高F1值0.950和ROC-AUC 0.989,同时在传播力前1%、5%、10%阈值下识别出高影响话题集群。对于无原生互动数据集,使用逻辑回归拟合的传播倾向得分作为扩散潜力的序数代理,而非直接互动度。结果表明,融合语义结构、传播感知排序与情感语言分析,可实现跨疫情的可扩展、可解释的比较分析。该框架支持面向监控的早期预警,能突出低频但高风险的叙事供分析师审查。

原文摘要 · Abstract (English)

Health misinformation circulating during pandemics can gain traction rapidly, creating harmful narratives that compete with public health guidance. Most topic-modelling pipelines treat engagement as an external outcome, limiting their ability to prioritise semantically coherent topics that are also rapidly diffusing. We introduce BERTopic-VP, a virality-prioritised topic-modelling framework that combines contextual embedding-based clustering (BERTopic) with a post hoc Virality Prioritisation (VP) layer. The pipeline is complemented by a two-stage hybrid misinformation detection module that fuses a supervised content-based classifier with an external verification signal derived from public-health knowledge bases. Applied to three benchmark datasets, COVID-19_FNIR, Monkeypox, and Constraint, the framework achieves strong classification performance, with F1 up to 0.950 and ROC-AUC up to 0.989, while identifying high-impact clusters under top 1%, 5%, and 10% VP thresholds. For datasets without native engagement metadata, prioritisation is based on a logistic propensity-to-spread score, used as an ordinal proxy for diffusion potential rather than a direct measure of engagement. The results show that integrating semantic structure, virality-aware ranking, and affective-linguistic profiling enables scalable and interpretable comparative analysis of misinformation across pandemics. The proposed framework supports monitoring-oriented early warning by surfacing low-volume but high-risk narratives for analyst review.

谣言检测主题建模社交媒体公共卫生

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。