发现并高效标注YouTube上成瘾症谣言,助力公共健康干预
MythTriage: Scalable Detection of Opioid Use Disorder Myths on a Video-Sharing Platform
- 用轻量模型初筛+大模型精审的分级标注流程
- 覆盖2.9万搜索结果与34.3万推荐视频,识别出8个普遍误区
- 标注成本降低76%,适合平台监管与公共卫生研究
了解网络健康信息中错误认知的传播情况,有助于制定公共健康政策。然而,对高风险但研究不足的话题如阿片类药物使用障碍(OUD)——美国主要死因之一——进行大规模误信息测量仍具挑战。本文首次对YouTube上与OUD相关的谣言进行了大规模研究。通过临床专家验证,确认了8个广泛存在的谬误,并发布了专家标注的视频数据集。为实现规模化标注,提出MythTriage高效分诊流程:使用轻量模型处理常规案例,将复杂样本转交高性能但昂贵的大语言模型(LLM)处理。该方法在宏平均F1得分达到0.86的同时,预计可比人工标注和全量LLM标注减少超过76%的标注时间和费用。通过对2.9千个搜索结果和34.3万条推荐内容的分析,揭示了谣言在YouTube上的传播机制,为公共健康干预和平台治理提供可操作洞见。
原文摘要 · Abstract (English)
Understanding the prevalence of misinformation in health topics online can inform public health policies and interventions. However, measuring such misinformation at scale remains a challenge, particularly for high-stakes but understudied topics like opioid-use disorder (OUD)--a leading cause of death in the U.S. We present the first large-scale study of OUD-related myths on YouTube, a widely-used platform for health information. With clinical experts, we validate 8 pervasive myths and release an expert-labeled video dataset. To scale labeling, we introduce MythTriage, an efficient triage pipeline that uses a lightweight model for routine cases and defers harder ones to a high-performing, but costlier, large language model (LLM). MythTriage achieves up to 0.86 macro F1-score while estimated to reduce annotation time and financial cost by over 76% compared to experts and full LLM labeling. We analyze 2.9K search results and 343K recommendations, uncovering how myths persist on YouTube and offering actionable insights for public health and platform moderation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。