arXiv:2504.00027cs.CLcs.AI2025-04

从Reddit提取阿片类药物使用信息,助力实时监测过量事件

Opioid Named Entity Recognition (ONER-2025) from Reddit

  • 构建含33万词元的阿片类药物实体标注数据集,覆盖8类用药信息
  • 基于Transformer模型实现97%准确率,较基线提升10.23%
  • 针对俚语和情绪化表达设计实时监测系统,适合公共卫生研究者

阿片类药物过量危机仍是美国重大公共卫生问题,造成大量死亡与社会成本。社交媒体平台如Reddit提供了关于阿片类药物使用公众认知、讨论与经历的海量非结构化数据。本研究利用自然语言处理技术,特别是阿片类药物命名实体识别(ONER-2025),从这些平台中提取可行动信息。研究做出四项贡献:第一,创建了来自Reddit的独特人工标注数据集,用户自述不同给药途径的阿片类药物使用经历,包含331,285个词元及8类主要阿片类药物实体;第二,详述标注流程与指南,并讨论了标注过程中遇到的挑战;第三,分析了阿片类药物讨论中的关键语言难题,包括俚语、歧义、碎片化句子和情绪化语言;第四,提出一个实时监测系统,用于处理来自社交媒体、医疗记录和急救服务的流式数据,以识别过量事件。通过11次实验的5折交叉验证,系统融合机器学习、深度学习与基于Transformer的语言模型,结合先进上下文嵌入技术以增强理解。基于Transformer的模型(bert-base-NER 和 roberta-base)在测试中达到97%准确率与F1分数,优于基线模型10.23%(随机森林=0.88)。

原文摘要 · Abstract (English)

The opioid overdose epidemic remains a critical public health crisis, particularly in the United States, leading to significant mortality and societal costs. Social media platforms like Reddit provide vast amounts of unstructured data that offer insights into public perceptions, discussions, and experiences related to opioid use. This study leverages Natural Language Processing (NLP), specifically Opioid Named Entity Recognition (ONER-2025), to extract actionable information from these platforms. Our research makes four key contributions. First, we created a unique, manually annotated dataset sourced from Reddit, where users share self-reported experiences of opioid use via different administration routes. This dataset contains 331,285 tokens and includes eight major opioid entity categories. Second, we detail our annotation process and guidelines while discussing the challenges of labeling the ONER-2025 dataset. Third, we analyze key linguistic challenges, including slang, ambiguity, fragmented sentences, and emotionally charged language, in opioid discussions. Fourth, we propose a real-time monitoring system to process streaming data from social media, healthcare records, and emergency services to identify overdose events. Using 5-fold cross-validation in 11 experiments, our system integrates machine learning, deep learning, and transformer-based language models with advanced contextual embeddings to enhance understanding. Our transformer-based models (bert-base-NER and roberta-base) achieved 97% accuracy and F1-score, outperforming baselines by 10.23% (RF=0.88).

命名实体识别公共卫生社交媒体分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。