arXiv:2605.26445cs.CL2026-05中稿 · IEEE International…

从Reddit收集药物使用真实经历,构建医疗级实体标注数据集

Curation and Extraction of Drug-Related Entities from Reddit Platform

论文配图:Curation and Extraction of Drug-Related Entities from Reddit Platform
图 1 · 摘自论文原文
  • 用毒理学专家标注6435篇Reddit帖子,提取药物、剂量、效果三类实体
  • BiomedBERT在药物识别上F1达0.843,大模型中Llama-3 70B表现最优
  • 首次系统捕捉患者自述用药体验,适合医学信息挖掘与社会媒体研究

医生主要通过临床过量案例了解非法药物,难以掌握真实使用情况。而药物使用者在社交媒体上分享第一手经验,提供关于剂量和药效的宝贵信息。为弥合这一差距,我们提出ReDose(REddit Drug DOSe and Effect)数据集,包含6,435篇关于物质使用的Reddit帖子。由一位持有行医执照的毒理学家主导标注训练集和测试集,两名医学生协助测试集标注,识别药物(DRUG)、剂量(DOSE)和效果(EFFECT)实体。我们基于6,267条标注对BERT类模型、大语言模型(LLM)及检索增强生成(RAG)模型进行了基准测试。BiomedBERT在药物识别上的F1得分为0.843,而Llama-3 70B表现优于GPT-4(F1分别为0.79和0.72)。效果提取仍具挑战,GPT-4的召回率仅为0.41。ReDose捕捉患者自述叙事,推动医疗数据从社交媒体中的自动提取。

原文摘要 · Abstract (English)

Physicians learn primarily about illicit drugs from clinical overdose cases, limiting their understanding of real-world usage. Meanwhile, drug users share first-hand experiences online, offering insights into dosage and effects of drugs. To bridge this gap, we introduce ReDose (REddit Drug DOSe and Effect), a dataset of 6,435 Reddit posts on substance use. A board-certified toxicologist primarily annotated both the training and test sets, while two medical science students contributed to the test set, labeling DRUG, DOSE, and EFFECT entities. We benchmarked 6,267 annotations using BERT-based, large language model (LLM)-based, and Retrieval-Augmented Generation (RAG) models. BiomedBERT achieved an F1-score of 0.843 for DRUG, while Llama-3 70B outperformed GPT-4 (F1 = 0.79 vs. 0.72). EFFECT extraction remains challenging, with GPT-4 achieving a recall of 0.41. ReDose captures patient-curated narratives to advance medical data extraction from social media.

药物实体抽取社交媒体医疗大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。